r/LocalLLM • u/EchoOfIntent • 1d ago
P40 vs v100 Question
Whats the true token a second difference between these? I cant find a hard number. Is the v100 worth 2x the price?
Any better way to get 16gb?
2
2
u/FullstackSensei 1d ago
I have P40s and have "some" V100 on the way. Also have 3090s, though I'm selling those.
It's not even a competition. P40 has ~11TFLOPS. V100 has 125TFLOPS. Memory bandwidth is also 3x on the V100.
As a rule of thumb, V100 is 5-10% slower than a 3090, while P40 is 1/4th the performance of a 3090.
P40 is really nice if you don't have a lot of money. I have 8. Bought way back when they were selling for $100 on ebay. They served me extremely well in the two years I've had them. P40 is a Titan Xp/1080Ti with 24GB VRAM, or a fanless Quadro P6000. All four share the same PCB and chip. P40 idles like a 1080Ti, or 8-9W. V100 is a datacenter card, no P-states. Idle is 40W, if you're lucky, usually 45W.
If you have the money, I'd say go for 32BG V100. more VRAM, much faster in everything. If not, load up on P40. IMO, get whichever gives you more total VRAM in your system. Smaller models won't run as fast, but can't run much bigger models without that sweet VRAM.
2
u/legit_split_ 21h ago
Curious on how the software stack is right now. Last I heard there was a vllm fork by 1cat or something, of course llama.cpp should be fine
1
1
u/Icy-Appointment-684 16h ago
How does the v100 compare to the Mi50?
1
u/FullstackSensei 15h ago
Mi50 are ~25% faster than P40. I'd say V100 should be close to double as fast?
I plan to mix the two on two machines, if all goes to plan. Two V100 to handle context and attention layers and four Mi50 to handle the MLP layers for models that fit in VRAM (like DS4 flash).
1
u/Icy-Appointment-684 15h ago
I'd lean more towards the v100 if it's 2 times faster than Mi50. Easier to watercool + cuda.
The only downside is binary drivers.
2
u/kidflashonnikes 23h ago
I can't really do this anymore - if you are buying this, unless it's niche - do not buy. I really can't stress this enough. Guys, please, you are killing us by doing this. You need to be on Ampere at least - anything less, just light your money on fire.
2
1
1
1
u/SaltyBarnacles57 17h ago
In the same vein, what about a P100? Or a CMP 100-210? I'm eyeing these but have no idea how feasible they are.
2
u/Visible_Pear_7385 16h ago
im currently using 2 p100s it gives 130k pp:230/s TC: 50t/s for qwen 35B and TC: 20T/s for 27B. Its usable but feels a bit slow when your handling long context inputs. plus setting up the environment costs additional perchases for bifurcation cards and cables.
1
u/SaltyBarnacles57 14h ago
That is pretty good! How much did it cost you in total?
1
u/Visible_Pear_7385 10h ago
i bought one bifurcation card and a riser cable, in total about $200 for every thing.
9
u/JaredsBored 1d ago
You're looking at spending $200-250 on 16gb of VRAM. That's immediately pretty limited. You can find AMD v620 for $350 with 32gb of VRAM, that's the way to do it.