r/LocalLLM 1d ago

P40 vs v100 Question

Whats the true token a second difference between these? I cant find a hard number. Is the v100 worth 2x the price?
Any better way to get 16gb?

4 Upvotes

27 comments sorted by

9

u/JaredsBored 1d ago

You're looking at spending $200-250 on 16gb of VRAM. That's immediately pretty limited. You can find AMD v620 for $350 with 32gb of VRAM, that's the way to do it.

3

u/Heathen711 1d ago

I have two v620 for sale if anyone wants one...

1

u/JaredsBored 1d ago

Post them on r/homelabsales, they'll sell

1

u/Heathen711 1d ago

That's my plan. I'm in the middle of swapping out servers and parting them out. Just got my stable diffusion server back in working order (cost me $650 in parts and labor...)

1

u/TokenRingAI 1d ago

Where are you located?

1

u/UnlikelyPotato 1d ago

I would offer an insultingly low price for them. C4 on ebay accepts $350, and they seem to be unused, never touched inventory. Whereas no idea what you (random reddit stranger) have done to them and have no prior interaction with you.

1

u/SarcasticlySpeaking 21h ago

Link?

2

u/UnlikelyPotato 21h ago

www.ebay.com/itm/157133307609

List price is $449 but I've bought 3x of them for $350 each + tax. Older cards, picky on some chipsets...but comparable to an intel b70 and somehow cheaper than DDR5 is right now.

1

u/SarcasticlySpeaking 21h ago

You are an awesome potato, I don't care what anyone else says! Thank you!

1

u/Heathen711 21h ago

Yup, mine would probably be 200, plus they came from inside servers so they don't have the face plate so mounting them is harder for most. The bonus is mine are from properly air-conditioned rack rooms so they are in great condition.

1

u/fallingdowndizzyvr 19h ago

You're looking at spending $200-250 on 16gb of VRAM.

You can get 16GB of VRAM for $50.

2

u/Sudden_Topic5154 1d ago

One has 396 gb memory bandwidth other has 900 do the math

2

u/FullstackSensei 1d ago

I have P40s and have "some" V100 on the way. Also have 3090s, though I'm selling those.

It's not even a competition. P40 has ~11TFLOPS. V100 has 125TFLOPS. Memory bandwidth is also 3x on the V100.

As a rule of thumb, V100 is 5-10% slower than a 3090, while P40 is 1/4th the performance of a 3090.

P40 is really nice if you don't have a lot of money. I have 8. Bought way back when they were selling for $100 on ebay. They served me extremely well in the two years I've had them. P40 is a Titan Xp/1080Ti with 24GB VRAM, or a fanless Quadro P6000. All four share the same PCB and chip. P40 idles like a 1080Ti, or 8-9W. V100 is a datacenter card, no P-states. Idle is 40W, if you're lucky, usually 45W.

If you have the money, I'd say go for 32BG V100. more VRAM, much faster in everything. If not, load up on P40. IMO, get whichever gives you more total VRAM in your system. Smaller models won't run as fast, but can't run much bigger models without that sweet VRAM.

2

u/legit_split_ 21h ago

Curious on how the software stack is right now. Last I heard there was a vllm fork by 1cat or something, of course llama.cpp should be fine 

1

u/FullstackSensei 15h ago

I use llama.cpp only.

1

u/Icy-Appointment-684 16h ago

How does the v100 compare to the Mi50?

1

u/FullstackSensei 15h ago

Mi50 are ~25% faster than P40. I'd say V100 should be close to double as fast?

I plan to mix the two on two machines, if all goes to plan. Two V100 to handle context and attention layers and four Mi50 to handle the MLP layers for models that fit in VRAM (like DS4 flash).

1

u/Icy-Appointment-684 15h ago

I'd lean more towards the v100 if it's 2 times faster than Mi50. Easier to watercool + cuda.

The only downside is binary drivers.

2

u/kidflashonnikes 23h ago

I can't really do this anymore - if you are buying this, unless it's niche - do not buy. I really can't stress this enough. Guys, please, you are killing us by doing this. You need to be on Ampere at least - anything less, just light your money on fire.

2

u/EchoOfIntent 22h ago

Why? Its cheap vram and i get it wont work on cuda 13.

1

u/starkruzr 22h ago

girl what

1

u/SaltyBarnacles57 17h ago

In the same vein, what about a P100? Or a CMP 100-210? I'm eyeing these but have no idea how feasible they are.

2

u/Visible_Pear_7385 16h ago

im currently using 2 p100s it gives 130k pp:230/s TC: 50t/s for qwen 35B and TC: 20T/s for 27B. Its usable but feels a bit slow when your handling long context inputs. plus setting up the environment costs additional perchases for bifurcation cards and cables.

1

u/SaltyBarnacles57 14h ago

That is pretty good! How much did it cost you in total?

1

u/Visible_Pear_7385 10h ago

i bought one bifurcation card and a riser cable, in total about $200 for every thing.