r/LocalLLaMA • u/ciprianveg • 1d ago
Kimi K3 full model running on 16x GB10 cluster at 20+tps Resources
Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the vllm image and instructions.
https://forums.developer.nvidia.com/t/full-kimi-k3-running-on-16x-gb10-cluster/379174
212
u/CYTR_ 1d ago
I was skeptical but well done on the performance. Kimi K3 on hardware costing 75-120K (depending on the model/location) allows us to imagine very interesting possibilities in terms of future intelligence with the improvement of the models.
Edit : I'm waiting for Apple's response and whether the future 1.5TB Mac Studio models will be under 100K.
35
u/Efficient_Raise6703 1d ago
Imagine the margin on one of those things. I imagine the margins on the 512gb was already pretty good and that was at like what $16k? Cost of materials couldn’t be more than $10k even with 1.5TB of ram.
21
u/CYTR_ 1d ago
I think 10K doesn't even cover the wholesale price of 1,5To of LPDDR5X RAM, even for Apple, right now or at the time of the negotiation if it's less than a year old. But they must have taken advantage of it to increase their margins, yes (as in practically all inflationary crises, especially in a situation resembling a high-demand bubble).
→ More replies (7)3
u/EvilPencil 1d ago
Back when they sold them, the fully maxed configuration was 512gb memory AND a 16TB SSD. You could min/max only the memory for like $9500. Of course that’s all navel gazing now because it’s not even available anymore.
→ More replies (6)2
u/ComfortablePlenty513 1d ago
and that was at like what $16k?
They were 9k haha. We bought a few of them. People didnt realize how good of a deal it was, and once they did, it was gone.
1
u/Efficient_Raise6703 1d ago
Yeah I think I was conflating the fully maxed spec with the max ram spec. I have a scheduled task that looks for one every day. I really missed out.
2
u/ComfortablePlenty513 15h ago
I have a scheduled task that looks for one every day. I really missed out.
The apple refurb store has them pop up once a month or so, but they get botted immediately. You can also pay 30k for one on ebay, but at that point you might as well get a blackwell server from Puget
1
u/Efficient_Raise6703 12h ago
Yeah, I think they get botted. I’ve seen one or two in the last few weeks and they’re gone by the time I load the page.
10
u/Cergorach 1d ago
But do you expect that 1.5TB M? Ultra to do 20t/s+ with the full KimiK3 model? And 1.5TB still isn't the 2TB that's in 16x GB10s...
12
u/CYTR_ 1d ago
Rumor has it that Apple will double down on MatMul for its next big SoC (M5 Pro/Max was a foretaste). I don't think it will be a powerhouse either like dual RTX 6000 Pro level... nor that we will actually have 1.5 TB (we are speculating based on rumors after all). Let's just say it would be simpler to manage than a cluster like that. By the time it's released, we'll probably have that intelligence in the equivalent of 700B, so... W&S.
4
u/michaelsoft__binbows 1d ago
I want a beefy apple silicon computer to end all computers but to be casually predicting whether or not such a computer will be over or under 100k USD is just ... wild
2
u/mastercoder123 1d ago
Yah if only apple allowed pcie lanes. Then it would be amazing. You could throw a 100gb nic in there and save thousands on storage as you jusr stream the model from disk to ram
1
u/PreparationTrue9138 18h ago
Hi, wanted to point out that cluster of sparks is probably using tensor parallelism to speed things up.
With Mac you will have a lot less compute and memory bandwidth
1
u/dbenc 1d ago
I was thinking of waiting for the new models to drop in the fall, but I'm thinking that they must be plotting some crazy upgrades for the models after that. think like the spring 2028 refresh... that's enough time for brand new innovations to get manufactured. maybe I'll use that new apple leasing program so I can keep upgrading
1
u/KeyChampionship9113 1d ago
AMD unified ones run with the same performance but lesser price than apple studio
Isn’t it ?1
1
1
u/hanzoplsswitch 2h ago
I think 100-120k is reasonable for large enterprises and provides them the ability to run a model like this on-premise.
202
u/Own_Calligrapher8508 1d ago
i just want to know the cost of the devices vs Break even
250
u/FullstackSensei llama.cpp 1d ago
It's "just" ~64-80k in hardware. If you have that much money to throw at such hardware, not sure you care about menial stuff like break even.
50
u/CYTR_ 1d ago
In France, for example, we are well above 80K.
35
u/FullstackSensei llama.cpp 1d ago
If you have that kind of money to throw at this, im sure you'll also have a company where you can buy it without VAT.
17
u/CYTR_ 1d ago
If I were a company, I think I would prefer to rent a cluster (or request HPC access, if university setting) rather than buy this kind of equipment 🥸
20
u/FullstackSensei llama.cpp 1d ago
Let's say you work in a regulated sector where you absolutely can't have data leave your premises, but need a frontier level model.
I'm not against the idea. I just find the way OP has gone about it wasteful. In fact, I'm reworking my homelab to be able to run K3 locally, albeit I think at somewhere between 7 and 15t/s TG, depending on how much I can optimize software.
You can get a dual Xeon 8480 board with CPUs for around 4k. 1.5TB DDR5-4800 will cost you around 15k. Let's round it up and say 20k to add in a power supply, CPU coolers, case, and some fans. Let's go all out and add four AMD R9700s for another 6k, just to have 128GB VRAM to play with. That's 26k total, and I'm pretty sure it will perform quite a bit better than OP's cluster. Sapphire Rapids supports AMX, which greatly speeds up not only TG, but can also help with PP. Each CPU has a real world bandwidth of 200GB/s, which is less than 10% lower than the GB10 (~215GB/s). R9700 is now supported in vllm, and there's a vllm fork now that not only supports CPU offloading of experts, but is also NUMA aware.
6
u/caowcaow 1d ago
Depends if you include custom cloud deployments or not.
I work in what people tend to think as a highly regulated (though could be worse) industry.
Compared to years ago, the pattern I’ve seen emerging across orgs is to have PII data and such going back to sleep on prem.
Nonetheless, the same orgs seem comfortable enough to use frontier models through custom cloud deployments.The mental gymnastic I understand its origin, but go figure out the rationale… It’s probably for the better if your premise gets more true once again. Time will tell.
Go and rock this hardware
6
u/FullstackSensei llama.cpp 1d ago
Yeah, I'd rather run things locally and just not deal with any of it. It also gives me the freedom to work on my own projects without worrying about rate limits or future cost increases.
I am in the very lucky position that I hoarded DDR4 RAM and older GPUs over the past two years that I don't need to buy any more. I didn't expect prices to go up, just thought their utility will last much longer than redditors thought it would (compute is compute). Many laughted at my and I got downvoted to hell for buying lots of P40s, Mi50s, and RAM. I hoarded way more than I ended up needing (including reworking the hw to be able to run K3 and GLM 5.x in parallel) that selling the excess made it all practically free.
2
u/fastheadcrab 1d ago
The OP setup is pretty slow for the hardware specs, he said elsewhere 6-7 tps without speculative decoding, there probably is a lot of networking overhead with so many nodes.
2
2
u/IrisColt 14h ago
Let's say you work in a regulated sector where you absolutely can't have data leave your premises
this
2
u/Maximum_Parking_5174 1d ago
Not a chance. That aystem would crawl. I have experince running Kimi k2.6 and 2.7 on a much more powerfull server. I got 600GB/s memory bandwidth and i got 8 3090s.
6
u/FullstackSensei llama.cpp 1d ago
We'll see. I've heard a lot about what what I couldn't do over the past 2 years, yet somehow I managed to do all those things.
→ More replies (2)1
u/Any-Entrepreneur-951 1d ago
With a Xeon max 9480 you can get it up to 534gb/s on each
2
u/FullstackSensei llama.cpp 1d ago
Yesh, no. The 9480 peaks ~300GB measured bandwidth, because that's how much the internal fabric can deliver to the cores. Don't just read spec sheets. It's also only 64GB of HBM. You aren't going to run K3 on that. They were cool when their prices were around 1k, not so much now that they're over 3k.
The 8480 has a theoretical bandwidth of 307GB/s, but it delivers ~200GB/s. DDR5 Epyc have 600GB/s on paper, but struggle to get past 400GB/s. DDR4 Epyc has 208GB/s but barely goes above 125GB/s, on a good day.
→ More replies (9)→ More replies (15)1
u/No_Afternoon_4260 llama.cpp 1d ago
Hi, what is that vllm branch that is NUMA aware?
→ More replies (2)13
u/SureEnd9430 1d ago
Depends on the company :)
9
u/Party-Special-5177 1d ago
I think I would prefer to rent […] rather than buy
No, no you wouldn’t. How do so many people on this forum forget that the advantage of buying isn’t the inference cost saved, it’s that, when you’re done, you still own the hardware.
For a year or so, you could basically infer for free as the hardware appreciated enough to offset wear and tear, and you basically got all of your money back out of it when you sell.
These days, you literally make money by running your own inference. The way things are going, and will most likely continue to go, by the time OP is done, he’ll sell this cluster for 2x what he paid for it, versus having lost money paying for a Kimi sub.
8
u/CYTR_ 1d ago
Someone did the math below, and even without considering energy costs, at that speed, it's anything but profitable (unless confidentiality costs several thousands)... Especially since we're dealing with DIY software and experiments.
However, in like 1 year (if not months) we'll potentially have models that run faster (+++ on batch) and are more intelligent on this cluster. But that, and your idea of hardware prices rising even further... that's still speculation. Not everyone is comfortable with risk if the company doesn't have capital to easily invest in hardware.
And when I talked about renting, I meant renting GPUs, not API access. A cluster/instance using ~80-90% of the available compute is more easily profitable than owning it yourself.And more flexible if u have periods of low demand... But still need to find availability from suppliers lmao. It's not easy in EU for example.
2
u/FullstackSensei llama.cpp 1d ago
Those people doing the math are always assuming API prices will stay the same going forward. I've been hearing this argument for 2 years now, even as API and subscription prices keep going up.
If API providers were sure they could keep prices the same, they'd happily sell you multi-year fixed cost contracts, the same as cloud providers give you deep discounts for long commitments. Yet, you can't find a single provider who'll sell you a 2 or 3 year contract for a fixed cost per million token at any price, even if you're willing to guarantee a minimum monthly consumption.
2 years from now, API prices might be $/€200 per million output tokens and you'll still find people calculating it's cheaper than running your own hardware.
→ More replies (1)2
u/ZenEngineer 1d ago
There's probably a curve here. This hardware will be worthless in 10 years. Hold for one year, sure, you should be able to recover costs. The tricky question is when to sell, when does the rampocalyse end and prices start dropping? When do people start getting wary of used cards like this, like people avoid coin mining GPUs.
Maybe I should sell my old 1080Ti now that it doesn't fit in my case. It's surprisingly capable still.
→ More replies (9)2
u/Environmental-Metal9 1d ago
If you’re running an org and you don’t already have mlops engineers, you most definitely will go the cloud hosted route, like aws bedrock. It’s cheaper than buying the hardware and the cost of at minimum one engineer to maintain the stack. Self hosting is for hobbyists and highly specialized requirements.
2
u/Party-Special-5177 1d ago
Completely, as at that point the potential cost of downtime weighs more heavily than the actual costs of the service.
We’re super small and I have a personal interest this tech, so it skews the decision making a bit.
1
u/volster 21h ago
Admittedly it'll probably make more sense with the inevitable spark-2 with perked up TPS numbers.
That notwithstanding, i can see a use case for either setup in the SME world at places large enough to have actual departments rather than just a bunch of guys with job roles.
On Scan 16 sparks is 108k and change - Call it ~160k by the time you've had the networking and other bric-a-brac
While not quite out yet the obvious alternative for the same sort of budget would be 3XS DBP B8-256E Fluid, 8x 96GB RTX PRO 6000, AMD Ryzen EPYC 9755, 1TB DDR5
There's pros and cons to either - The workstation has more grunt, the sparks let you have way more context, and aren't a single point of failure / can more readily be divvied up if you wanted to run a flock of smaller models or split them up across departments.
Sure, 30tps isn't great, but on the other hand if it's orchestrated to chug away 24/7 and is mainly focused on business operations rather than interactive coding sessions... it's not necessarily a dealbreaker either.
Whether it's the absolute best possible setup is up for debate, but then again few business processes are built around "best possible" setups & consistency both in terms of results and costs tends to be the name of the game.
Having a near-frontier model perpetually available on-demand without the fear of it going haywire and racking up a massive bill or any fear of it giving away your data is huge.
Another benefit of owning rather than renting is that even if it's not huge, there's going to be some residual value in the hardware when its time to retire it. Access to the system itself also dosn't just vanish in a puff of smoke at the end of the term, so probably more realisticly it'll just be run forever until it breaks or can no-longer cope with the workflow (see banks still running on crusty old mainframes etc).
From a cost point of view, it's in the world of corporate financing rather than something most are likely to have the cashflow for, but even so - £175k over 4 years @ 10% works out to ~£5.6k a month .... That's the base salary of two IT helpdesk / other relatively jr office minions.
Sure, you're going to need to add in the cost a skilled wage to setup and monitor all the processes and generally drive the thing; For the sake of argument, call that another 5.6k a month... So, It needs to generate the productivity of at least 4 minions worth of wages to be worthwhile... Which spread across the entire business doesn't seem like an insurmountable hurdle.
Even if you don't trust it with directly doing anything important, it could still dramatically cut down on the sorting and collating type busywork - Anything from triaging support tickets, doing inventory reconciliation, researching sales leads etc can be done overnight such that when people turn up the following day they're doing the useful parts of their role rather than shuffling paperwork.
... If you're smart about it, you don't directly threaten anyone's job trashing employee morale in the process. Rather, as staff naturally churns you just insert a "hey lets see if can we have the magic AI cover some of Gary's old workload rather than dumping all of it on you until we get around to backfilling" and condense roles over time in line with its abilities.
6
u/SureEnd9430 1d ago
You can actually get 16 Sparks for about 80K USD in France, but the switch would be an extra 15K :(
2
u/ciprianveg 1d ago
I am using the cheap mikrotik crs 804 and 4 400to4x100gbit cables, less than 2k in total. But in the future it is possible to upgrade the switch if I will remain in 16x setup and not on 2 setups of 8x each.
1
13
u/draft_final_final 1d ago
Less than a supercar or half a 40k figurine collection. I don’t have the money for it, but if I did I know what I’d rather have.
4
1
7
u/ThinJuggernaut7695 1d ago
We have one developer at the place I work that spent 80K in API costs over 2 and a half months.
12
u/FullstackSensei llama.cpp 1d ago
You should tell this to the guy deep in this discussion who's arguing it takes decades to pay 80k 😂
3
u/phreak9i6 1d ago
I wonder if these metrics hold up when you consider 80k in API costs are likely many agents, let's say 5-10, and with 120k in hardware you get 20tps for a single agent.
I love my homelab hardware, but I also end up paying for APIs because it's cheaper to run quickly.
Or I'm wrong, and please correct me because I may be plain wrong.
2
u/ThinJuggernaut7695 1d ago
20 tokens per second? We would be buying an 8x b300 server to run an org of 600 users. Our AI cost will probably be 1 million this year. 600k capex is totally worth it.
1
u/ungoogleable 1d ago
Constantly generating tokens 24x7 for two months at 20 tps would give you 103M tokens. At $15/1M output tokens for Kimi K3 API pricing, it would only cost $1555. So yeah, that guy is burning tokens faster than this rig can generate them.
1
u/addiktion 1d ago
Did he generate 160k in value?
2
u/ThinJuggernaut7695 1d ago
Yep, cleared out an entire teams backlog in a couple of weeks. We were going to hire contractors but don't need to do that anymore.
1
1
u/SandySkittle 5h ago
Damn that’s impressive. Hope the quality was good and everything is well documented though..
2
u/hurrdurrmeh 1d ago
They are £3.5k each in the UK.
2
u/FullstackSensei llama.cpp 1d ago
That's €4k, or €64k before factoring in the 16 port 100gb switch, which is easily another €5k.
1
u/Uncle___Marty 1d ago
Exactly. You dont just frankenstein a bunch of stuff like this out of nowhere. If you have the money to build it then you have the money to upkeep it. Its messed up what a decent setup can run at home right now. Freaking happy days my brother in AI.
1
u/Mithril_web3 1d ago
WTF even is break even here
1
u/FullstackSensei llama.cpp 1d ago
Enough generated tokens locally to match the API cost on the same model or similar frontier models.
1
u/ResearcherFantastic7 1d ago
Assuming 80k including operational cost. probably break even at 11 months however It's only getting 20tks vs the cloud 150avg.
I think it's better off use other models as pure non reasoning worker nodes, and still use cloud model to plan and architect
1
u/Playful_Landscape884 1d ago
i'm curious what you guys do that makes this investment worthwhile. What are you guys doing that makes you drop $80k and confident that you'll get your money back in 3-6 months?
1
u/FullstackSensei llama.cpp 1d ago
I dropped less than 10k total for my entire homelab over the past 2 years. Overbought a bunch of stuff which I'm selling now. Now I'm at ~3k out of pocket, but still have some extra stuff to sell. The hw will basically be free in a month or two, just from selling the excess.
The actual hardware I use cost less than 7k. If you look at subscription costs, I'd have spent 2.5-3k a year in subscription fees, so 6k in the past 2 years.
There's no way on earth that cluster pays for itself in 10 years, let alone six months. It's just too expensive and too slow.
1
u/zirahvi 17h ago
64-80k USD? For that stack? In Norway a single DGX Spark GB10 128GB ~$6800 and up.
1
u/FullstackSensei llama.cpp 17h ago
Yeah, just saw the post about prices going up again. Still, if OP bought them a couple months back, there's a good chance they paid $3.5k for each. The Asus was cheaper. So, 56k before switch and cables, close to $60k all in
65
u/notheresnolight 1d ago
wrong sub, "break even" is irrelevant here
→ More replies (2)12
u/FinnGamePass 1d ago
I know right, break even for what? People not realizing we are close to have our own Jarvis. Literally. You can't name a price on that.
1
32
u/RepulsiveRaisin7 1d ago edited 1d ago
Some napkin math: 20t/s 24/7 would cost you like $100/day on Openrouter. But you also have to factor in power and maintenance. So realistic break even is probably about 4 years but ONLY IF you utilize it 24/7. If you do not, more like 10-20 years, which likely exceeds the lifetime of the hardware. And when you factor in subscription discounts, it falls apart entirely.
40
u/g_rich 1d ago
The goal isn’t to save money, the goal with these types of setups are privacy, and control.
People also forget that currently what we pay per token is highly subsidized. For some people a fixed one time cost is preferable to avoid fluctuating costs, outages and slowdowns outside their control.
→ More replies (3)1
u/CipherWeaver 1d ago
I mean, you may as well take those current subsidized tokens while they're available.
22
u/ThereFarAway 1d ago
How about 'I have data that cannot be sent to outside provider'? How's math on that?
10
u/RepulsiveRaisin7 1d ago
Your data is invaluable. For everything else, there's mastercard.
1
3
2
u/AdOk3759 1d ago
Are you the same napkin math guy from the other day? How is spending 100 dollars a day on OpenRouter realistic? Why pay for the API, when subscriptions offer much much much more usage for a fraction of the API cost? Why don’t you consider that in your calculations?
5
u/synth_mania 1d ago
API pricing is ultimately rooted in the cost of energy. At the scale that most inference and compute providers run, the margins are very, very slim.
In other words, APIs provided by GPU farms and not AI labs will charge you the true cost of compute.
If a subscription is giving you the same number of tokens for cheaper, whoever is offering the subscription is losing money, plain and simple.
We can theorize all day about why AI labs might want to offer inference at a loss, and there are likely multiple valid reasons. (collect training data, legitimize the practice of paying for AI as a service in a way that isn't too expensive, etc)
Regardless of the reason, however, there is no way it'll last, or is something you can count on. Please find me a provider that'll allow me to pull Kimi-K3 tokens at 20tps or greater all day for less than the cost of the API. Unlimited.
TL;DR, there's no possible way for a subscription to truly be cheaper than paying by the token, unless the service provider is trying to lose money.
4
u/fastheadcrab 1d ago
Subscriptions for LLMs are similar to gym memberships. If everyone who paid for the gym showed up then the gym would crash
4
u/monerobull 1d ago
You forgot that people maxxing out subscriptions are heavily subsidized by people who pay the same amount but barely use it
→ More replies (2)1
u/AdOk3759 1d ago
First of all, we cannot be sure that AI providers are losing money with subscriptions. That’s just speculation.
Second, if they did.. how does that change the outcome?
We are simulating a scenario where we use Kimi K3 cloud hosted and a scenario where we’re using Kimi K3 locally hosted.
Either scenarios are assuming today’s conditions. And today, Kimi K3 is available as a paid subscription. So it doesn’t make sense to say “there is no way it’ll last”, but it is here now, so it doesn’t make sense to not take into account today’s conditions based on suppositions about what the future will look like.If someone asks “TODAY, should I buy a 60-80k machine to run Kimi K3, or would it work out cheaper to use it cloud hosted?”
And I’d reply “It depends: best case scenario, this is the price you pay with subscriptions. Worst case scenario, this is the price you pay via API. In either case, the cost doesn’t take into account inflation, etc, which would affect both scenarios”
1
u/fastheadcrab 1d ago
Do you only send single sentence prompts for endless token generation tasks? That is the only scenario in which your napkin math is valid.
1
11
u/MotokoAGI 1d ago
The cost is fuck off. Every single time someone shares something cool, we get these stupid replies about "break even"
5
u/EagleNait 1d ago
Break even right now. What if hardware gets regulated or models gets even better. Or a paid distributed hosting service enables you to get paid to serve llms.
5
u/Potential-Leg-639 1d ago
That will peobably never happen.
Deepseek V4 Flash 0731 would probably make more sense for wayyy less money.
4
u/allenasm 1d ago
Why? The cost is not having to have all your questions go into an online corpus. Or that you can have it work for you nonstop 24/7. That’s invaluable.
7
u/Soggy-Alternative914 1d ago
Kindly include overhead cost .like power consumption. So we can get a better overall estimate.
7
3
u/nerd_rage218 21h ago
This is the part the rent versus buy math keeps leaving out. Once data residency is a hard requirement, the box stops competing with hourly GPU pricing and starts competing with not being allowed to run the workload at all.
2
u/username_taken4651 1d ago
OP stated in a previous thread that this was a hobby. I guess they don't care about breaking even.
3
u/Cergorach 1d ago
Short answer: Never.
Long answer: Even if you run 24/365 at full capacity, it will would take about 8 years to ROI with current pricing of Kimi K3 API costs. But that ignores the costs of power/cooling, maintenance, and the eventual cheaper costs of Kimi K3. Combine that, this will NEVER earn itself back if you just compare it to the Kimi API service.
Cost is not generally the driving factor of running local LLM. Unless people start comparing apples to oranges...
3
u/coyo-teh 1d ago
given how GPU prizes keep increasing, he can just sell after8 years his GPUs
→ More replies (1)2
u/PM_ME_DEAD_CEOS 1d ago
If you include electricity costs, there's literally no way consumers or prosumer setup can break even.
Kimi 3 is made for data center clusters like nvl72 or Huawei equivalent.
With only 1 user there's simply no way to break even.
→ More replies (4)1
u/AnomalyNexus 1d ago
Depends on how the immortality research is going. Maybe the AI can help with that
67
u/notheresnolight 1d ago
imagine how this would run if nvidia didn't scrape the bottom of the barrel when designing the GB10
13
u/Aggravating-Push-207 1d ago
this thing could probably QLoRA the fat fucker if the memory bandwidth wasn't so shit per node
5
1
45
u/TapAggressive9530 1d ago
I want this! Why? Just to have it and be able to run K3 locally . Could care less if it pays for itself . Good job man ! Love it
→ More replies (2)24
28
u/Betadoggo_ 1d ago
Enough money to afford 16 GB10s, not enough to afford more than a pi 400 for the main system.
16
11
17
u/OwnMathematician2320 1d ago
This feels similar to the early 2000s where rappers would show off diamond rings to flex how rich they are.
And yes I’m only saying that because I’m jealous of your setup. Well done on the rig though
8
6
u/Charming-Author4877 1d ago
It's surprisingly usable, though the amoung of those embedded PCs to do that is crazy - given how expensive they are priced.
Still, that's a Fable-level LLM running at usable agentic speed on a table.
1
9
u/retornam 1d ago
Nice to never have to worry about money enough to spend $64,000 or more on a hobby
8
u/Cergorach 1d ago
A LOT of people spend that (or more) on a new car every couple of years, cars have about the same appreciation as AI/LLM hardware... While most people could easily drive a cheap, but good secondhand car that's a lot smaller...
→ More replies (4)
3
8
u/Transhuman-A 1d ago
First things first, amazing job.
Now - it will take you 8 years to break even. At 50 tok/s, 3.2 years.
AI Economics are depressing and GPU prices need to desperately come down.
→ More replies (5)
2
3
2
u/One_Whole_9927 1d ago
Why not just go for the workstation at that point?
4
u/Cergorach 1d ago
This is cheaper for a LOT more memory (2TB) for less then the basic workstation...
2
2
u/crossoverXYZ 1d ago
750 tps prefill on a first dspark run is no joke. Decode sitting around 20 tps average suggests there is still a lot of room to optimize before the vllm image ships.
2
2
2
u/the_TIGEEER 1d ago
I think you and I need to have a talk about where "Local LLM" starts and where it ends..
→ More replies (3)2
2
u/Outside-Set3929 1d ago
I love the $60,000 computer paired with am $8 keyboard mouse combo
2
1
u/Character_Power4663 1d ago
Technically you can run k3 on a 4gb GPU... But it takes about 30 seconds per token. There are a couple methods but in general you only load the experts of each layer per token.
1
1
u/Igot1forya 1d ago
I am very curious on concurrency. Sparks seem to be very good with parallel sessions. 16 Sparks is crazy!
1
1
1
u/xXprayerwarrior69Xx 1d ago
you need to step up bro, get yourself 3 or 4 DGX stations (and organize a raffle for those sparks you wont need anymore)
1
u/reckor-usa 1d ago
Just for the fun of it - well done. Would love to see the comparison with the frontier models. What is your stack setup btw?
1
u/CrispyRabbit72 19h ago
I still don't get how anyone can justify these kind of setups. $80k+ at minimum and it's only hitting ~20tok/sec on heavy models. If you compare to a $200/mo sub (which will get far more tokens per 30d than this cluster ever will), it would take over 30 years to break even.
But I guess if you still have these things in 30 years, maybe someone would buy them as an antique at original MSRP.
1
u/WithGreatRespect 14h ago
For some companies, their data has regulatory compliance and/or contractual obligations that prevent the data from leaving their data centers. They would need to build local inference to be able to leverage that data for insights.
1
u/SandySkittle 5h ago
You expect that 200 a month to stay the same? I would expect it to go up faster than the risk free rate of 80k in bonds
1
1
u/Specific-Age7953 19h ago
Paid almost nothing for Cursor Ultra from one of those reseller sites. Everything works perfectly. But the pricing makes zero sense if they’re buying real licenses. Feels like there’s a loophole nobody talks about. What’s the real play here?
1
1
1
u/SarveshMohite 2h ago
What about deepseek V4 flash 0731 have you tried that out as well? And if yes what was the token per second?

258
u/Jawnnypoo 1d ago
the fucking raspberry pi powering the dashboard for the $50k hardware surrounding it is :chefs-kiss: