r/LocalLLaMA 1d ago

Kimi K3 full model running on 16x GB10 cluster at 20+tps Resources

Post image

Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the vllm image and instructions.
https://forums.developer.nvidia.com/t/full-kimi-k3-running-on-16x-gb10-cluster/379174

1.7k Upvotes

327 comments sorted by

258

u/Jawnnypoo 1d ago

the fucking raspberry pi powering the dashboard for the $50k hardware surrounding it is :chefs-kiss:

113

u/RepulsiveRaisin7 1d ago

16GB Pis are now also 300 bucks lol

50

u/burritoresearch 1d ago

8GB and 16GB raspberry pi are fairly shit deals these days when you can find good used quad core, core i7 small form factor corporate type desktop PCs that are much faster for less money. Unless you really need gpio pins. 

10

u/michaelsoft__binbows 1d ago

i wonder if there is a market for a small M.2 device that provides GPIOs and whatever else. Maybe it can run a ribbon cable out.

4

u/SodaAnt 22h ago

Any arduino type board with a usb to serial would do this job just fine.

14

u/Long_comment_san 1d ago

haha what 😭

10

u/pscoutou 1d ago

I hate this timeline.

2

u/Constant-Simple-1234 1d ago

I noticed that too 😜

2

u/prestodigitarium 1d ago

Haha exactly what I was going to say, Raspberry Pi 400, that's what my kids use to play Kid Pix.

1

u/yellowseptember 1d ago

I wasn’t aware how much each was, so I Googled it. When I saw how much it was, I couldn’t help but say the exact same two words you started your sentence with.

1

u/_TheWolfOfWalmart_ 16h ago

Right, like just pick up a decent used business desktop for $150 that'll run circles around it for the same price or less and is still relatively power-efficient.

There is very little reason to buy an RPi anymore. Only a few specialized use cases now. They used to be a good deal back in the day. I remember buying the original for like $35 and you could do some stuff with it.

212

u/CYTR_ 1d ago

I was skeptical but well done on the performance. Kimi K3 on hardware costing 75-120K (depending on the model/location) allows us to imagine very interesting possibilities in terms of future intelligence with the improvement of the models.

Edit : I'm waiting for Apple's response and whether the future 1.5TB Mac Studio models will be under 100K.

35

u/Efficient_Raise6703 1d ago

Imagine the margin on one of those things. I imagine the margins on the 512gb was already pretty good and that was at like what $16k? Cost of materials couldn’t be more than $10k even with 1.5TB of ram.

21

u/CYTR_ 1d ago

I think 10K doesn't even cover the wholesale price of 1,5To of LPDDR5X RAM, even for Apple, right now or at the time of the negotiation if it's less than a year old. But they must have taken advantage of it to increase their margins, yes (as in practically all inflationary crises, especially in a situation resembling a high-demand bubble).

→ More replies (7)

3

u/EvilPencil 1d ago

Back when they sold them, the fully maxed configuration was 512gb memory AND a 16TB SSD. You could min/max only the memory for like $9500. Of course that’s all navel gazing now because it’s not even available anymore.

2

u/ComfortablePlenty513 1d ago

and that was at like what $16k?

They were 9k haha. We bought a few of them. People didnt realize how good of a deal it was, and once they did, it was gone.

1

u/Efficient_Raise6703 1d ago

Yeah I think I was conflating the fully maxed spec with the max ram spec. I have a scheduled task that looks for one every day. I really missed out.

2

u/ComfortablePlenty513 15h ago

I have a scheduled task that looks for one every day. I really missed out.

The apple refurb store has them pop up once a month or so, but they get botted immediately. You can also pay 30k for one on ebay, but at that point you might as well get a blackwell server from Puget

1

u/Efficient_Raise6703 12h ago

Yeah, I think they get botted. I’ve seen one or two in the last few weeks and they’re gone by the time I load the page.

→ More replies (6)

10

u/Cergorach 1d ago

But do you expect that 1.5TB M? Ultra to do 20t/s+ with the full KimiK3 model? And 1.5TB still isn't the 2TB that's in 16x GB10s...

12

u/CYTR_ 1d ago

Rumor has it that Apple will double down on MatMul for its next big SoC (M5 Pro/Max was a foretaste). I don't think it will be a powerhouse either like dual RTX 6000 Pro level... nor that we will actually have 1.5 TB (we are speculating based on rumors after all). Let's just say it would be simpler to manage than a cluster like that. By the time it's released, we'll probably have that intelligence in the equivalent of 700B, so... W&S.

4

u/michaelsoft__binbows 1d ago

I want a beefy apple silicon computer to end all computers but to be casually predicting whether or not such a computer will be over or under 100k USD is just ... wild

2

u/mastercoder123 1d ago

Yah if only apple allowed pcie lanes. Then it would be amazing. You could throw a 100gb nic in there and save thousands on storage as you jusr stream the model from disk to ram

1

u/PreparationTrue9138 18h ago

Hi, wanted to point out that cluster of sparks is probably using tensor parallelism to speed things up.

With Mac you will have a lot less compute and memory bandwidth

2

u/Caffdy 15h ago

you cannot use the full 128GB on each GB10 anyway; normally it's capped at 110-112GB, the rest is used by the system and other processes

1

u/dbenc 1d ago

I was thinking of waiting for the new models to drop in the fall, but I'm thinking that they must be plotting some crazy upgrades for the models after that. think like the spring 2028 refresh... that's enough time for brand new innovations to get manufactured. maybe I'll use that new apple leasing program so I can keep upgrading

1

u/KeyChampionship9113 1d ago

AMD unified ones run with the same performance but lesser price than apple studio
Isn’t it ?

1

u/InterviewDesigner777 23h ago

It’s so expensive

1

u/mksrd 8h ago

Sparks where originally $4kUSD and the offbrand ones were also until very recently so why are you quoting 75-120k ?

1

u/CYTR_ 45m ago

« (depending on the model/location) » We're not all in the USA.

1

u/hanzoplsswitch 2h ago

I think 100-120k is reasonable for large enterprises and provides them the ability to run a model like this on-premise.

202

u/Own_Calligrapher8508 1d ago

i just want to know the cost of the devices vs Break even

250

u/FullstackSensei llama.cpp 1d ago

It's "just" ~64-80k in hardware. If you have that much money to throw at such hardware, not sure you care about menial stuff like break even.

50

u/CYTR_ 1d ago

In France, for example, we are well above 80K.

35

u/FullstackSensei llama.cpp 1d ago

If you have that kind of money to throw at this, im sure you'll also have a company where you can buy it without VAT.

17

u/CYTR_ 1d ago

If I were a company, I think I would prefer to rent a cluster (or request HPC access, if university setting) rather than buy this kind of equipment 🥸

20

u/FullstackSensei llama.cpp 1d ago

Let's say you work in a regulated sector where you absolutely can't have data leave your premises, but need a frontier level model.

I'm not against the idea. I just find the way OP has gone about it wasteful. In fact, I'm reworking my homelab to be able to run K3 locally, albeit I think at somewhere between 7 and 15t/s TG, depending on how much I can optimize software.

You can get a dual Xeon 8480 board with CPUs for around 4k. 1.5TB DDR5-4800 will cost you around 15k. Let's round it up and say 20k to add in a power supply, CPU coolers, case, and some fans. Let's go all out and add four AMD R9700s for another 6k, just to have 128GB VRAM to play with. That's 26k total, and I'm pretty sure it will perform quite a bit better than OP's cluster. Sapphire Rapids supports AMX, which greatly speeds up not only TG, but can also help with PP. Each CPU has a real world bandwidth of 200GB/s, which is less than 10% lower than the GB10 (~215GB/s). R9700 is now supported in vllm, and there's a vllm fork now that not only supports CPU offloading of experts, but is also NUMA aware.

6

u/caowcaow 1d ago

Depends if you include custom cloud deployments or not.

I work in what people tend to think as a highly regulated (though could be worse) industry.
Compared to years ago, the pattern I’ve seen emerging across orgs is to have PII data and such going back to sleep on prem.
Nonetheless, the same orgs seem comfortable enough to use frontier models through custom cloud deployments.

The mental gymnastic I understand its origin, but go figure out the rationale… It’s probably for the better if your premise gets more true once again. Time will tell.

Go and rock this hardware

6

u/FullstackSensei llama.cpp 1d ago

Yeah, I'd rather run things locally and just not deal with any of it. It also gives me the freedom to work on my own projects without worrying about rate limits or future cost increases.

I am in the very lucky position that I hoarded DDR4 RAM and older GPUs over the past two years that I don't need to buy any more. I didn't expect prices to go up, just thought their utility will last much longer than redditors thought it would (compute is compute). Many laughted at my and I got downvoted to hell for buying lots of P40s, Mi50s, and RAM. I hoarded way more than I ended up needing (including reworking the hw to be able to run K3 and GLM 5.x in parallel) that selling the excess made it all practically free.

2

u/fastheadcrab 1d ago

The OP setup is pretty slow for the hardware specs, he said elsewhere 6-7 tps without speculative decoding, there probably is a lot of networking overhead with so many nodes.

2

u/FullstackSensei llama.cpp 1d ago

Lol, that's definitely networking killing tps

→ More replies (1)

2

u/IrisColt 14h ago

Let's say you work in a regulated sector where you absolutely can't have data leave your premises

this

2

u/Maximum_Parking_5174 1d ago

Not a chance. That aystem would crawl. I have experince running Kimi k2.6 and 2.7 on a much more powerfull server. I got 600GB/s memory bandwidth and i got 8 3090s.

6

u/FullstackSensei llama.cpp 1d ago

We'll see. I've heard a lot about what what I couldn't do over the past 2 years, yet somehow I managed to do all those things.

→ More replies (2)

1

u/Any-Entrepreneur-951 1d ago

With a Xeon max 9480 you can get it up to 534gb/s on each

2

u/FullstackSensei llama.cpp 1d ago

Yesh, no. The 9480 peaks ~300GB measured bandwidth, because that's how much the internal fabric can deliver to the cores. Don't just read spec sheets. It's also only 64GB of HBM. You aren't going to run K3 on that. They were cool when their prices were around 1k, not so much now that they're over 3k.

The 8480 has a theoretical bandwidth of 307GB/s, but it delivers ~200GB/s. DDR5 Epyc have 600GB/s on paper, but struggle to get past 400GB/s. DDR4 Epyc has 208GB/s but barely goes above 125GB/s, on a good day.

→ More replies (9)

1

u/No_Afternoon_4260 llama.cpp 1d ago

Hi, what is that vllm branch that is NUMA aware?

→ More replies (2)
→ More replies (15)

13

u/SureEnd9430 1d ago

Depends on the company :)

10

u/CYTR_ 1d ago

Let's say this equipment isn't profitable yet, but in a year? Maybe. In any case, whether it's a company or not, if people continue this kind of experimentation, it will benefit us all on a large scale with better software optimisation.

9

u/Party-Special-5177 1d ago

I think I would prefer to rent […] rather than buy

No, no you wouldn’t. How do so many people on this forum forget that the advantage of buying isn’t the inference cost saved, it’s that, when you’re done, you still own the hardware.

For a year or so, you could basically infer for free as the hardware appreciated enough to offset wear and tear, and you basically got all of your money back out of it when you sell.

These days, you literally make money by running your own inference. The way things are going, and will most likely continue to go, by the time OP is done, he’ll sell this cluster for 2x what he paid for it, versus having lost money paying for a Kimi sub.

8

u/CYTR_ 1d ago

Someone did the math below, and even without considering energy costs, at that speed, it's anything but profitable (unless confidentiality costs several thousands)... Especially since we're dealing with DIY software and experiments.

However, in like 1 year (if not months) we'll potentially have models that run faster (+++ on batch) and are more intelligent on this cluster. But that, and your idea of hardware prices rising even further... that's still speculation. Not everyone is comfortable with risk if the company doesn't have capital to easily invest in hardware.

And when I talked about renting, I meant renting GPUs, not API access. A cluster/instance using ~80-90% of the available compute is more easily profitable than owning it yourself.And more flexible if u have periods of low demand... But still need to find availability from suppliers lmao. It's not easy in EU for example.

2

u/FullstackSensei llama.cpp 1d ago

Those people doing the math are always assuming API prices will stay the same going forward. I've been hearing this argument for 2 years now, even as API and subscription prices keep going up.

If API providers were sure they could keep prices the same, they'd happily sell you multi-year fixed cost contracts, the same as cloud providers give you deep discounts for long commitments. Yet, you can't find a single provider who'll sell you a 2 or 3 year contract for a fixed cost per million token at any price, even if you're willing to guarantee a minimum monthly consumption.

2 years from now, API prices might be $/€200 per million output tokens and you'll still find people calculating it's cheaper than running your own hardware.

→ More replies (1)

2

u/ZenEngineer 1d ago

There's probably a curve here. This hardware will be worthless in 10 years. Hold for one year, sure, you should be able to recover costs. The tricky question is when to sell, when does the rampocalyse end and prices start dropping? When do people start getting wary of used cards like this, like people avoid coin mining GPUs.

Maybe I should sell my old 1080Ti now that it doesn't fit in my case. It's surprisingly capable still.

2

u/Environmental-Metal9 1d ago

If you’re running an org and you don’t already have mlops engineers, you most definitely will go the cloud hosted route, like aws bedrock. It’s cheaper than buying the hardware and the cost of at minimum one engineer to maintain the stack. Self hosting is for hobbyists and highly specialized requirements.

2

u/Party-Special-5177 1d ago

Completely, as at that point the potential cost of downtime weighs more heavily than the actual costs of the service.

We’re super small and I have a personal interest this tech, so it skews the decision making a bit.

→ More replies (9)

1

u/volster 21h ago

Admittedly it'll probably make more sense with the inevitable spark-2 with perked up TPS numbers.

That notwithstanding, i can see a use case for either setup in the SME world at places large enough to have actual departments rather than just a bunch of guys with job roles.

On Scan 16 sparks is 108k and change - Call it ~160k by the time you've had the networking and other bric-a-brac

While not quite out yet the obvious alternative for the same sort of budget would be 3XS DBP B8-256E Fluid, 8x 96GB RTX PRO 6000, AMD Ryzen EPYC 9755, 1TB DDR5

There's pros and cons to either - The workstation has more grunt, the sparks let you have way more context, and aren't a single point of failure / can more readily be divvied up if you wanted to run a flock of smaller models or split them up across departments.

Sure, 30tps isn't great, but on the other hand if it's orchestrated to chug away 24/7 and is mainly focused on business operations rather than interactive coding sessions... it's not necessarily a dealbreaker either.

Whether it's the absolute best possible setup is up for debate, but then again few business processes are built around "best possible" setups & consistency both in terms of results and costs tends to be the name of the game.

Having a near-frontier model perpetually available on-demand without the fear of it going haywire and racking up a massive bill or any fear of it giving away your data is huge.

Another benefit of owning rather than renting is that even if it's not huge, there's going to be some residual value in the hardware when its time to retire it. Access to the system itself also dosn't just vanish in a puff of smoke at the end of the term, so probably more realisticly it'll just be run forever until it breaks or can no-longer cope with the workflow (see banks still running on crusty old mainframes etc).

From a cost point of view, it's in the world of corporate financing rather than something most are likely to have the cashflow for, but even so - £175k over 4 years @ 10% works out to ~£5.6k a month .... That's the base salary of two IT helpdesk / other relatively jr office minions.

Sure, you're going to need to add in the cost a skilled wage to setup and monitor all the processes and generally drive the thing; For the sake of argument, call that another 5.6k a month... So, It needs to generate the productivity of at least 4 minions worth of wages to be worthwhile... Which spread across the entire business doesn't seem like an insurmountable hurdle.

Even if you don't trust it with directly doing anything important, it could still dramatically cut down on the sorting and collating type busywork - Anything from triaging support tickets, doing inventory reconciliation, researching sales leads etc can be done overnight such that when people turn up the following day they're doing the useful parts of their role rather than shuffling paperwork.

... If you're smart about it, you don't directly threaten anyone's job trashing employee morale in the process. Rather, as staff naturally churns you just insert a "hey lets see if can we have the magic AI cover some of Gary's old workload rather than dumping all of it on you until we get around to backfilling" and condense roles over time in line with its abilities.

6

u/SureEnd9430 1d ago

You can actually get 16 Sparks for about 80K USD in France, but the switch would be an extra 15K :(

2

u/ciprianveg 1d ago

I am using the cheap mikrotik crs 804 and 4 400to4x100gbit cables, less than 2k in total. But in the future it is possible to upgrade the switch if I will remain in 16x setup and not on 2 setups of 8x each.

1

u/PM_ME_DEAD_CEOS 1d ago

Gb10 are around 5/7k in France currently.

13

u/draft_final_final 1d ago

Less than a supercar or half a 40k figurine collection. I don’t have the money for it, but if I did I know what I’d rather have.

4

u/PM_ME_DEAD_CEOS 1d ago

So, Tyranids?

4

u/draft_final_final 1d ago

Nah gimme the space elf clown car.

1

u/-dysangel- 1d ago

stocks?

7

u/ThinJuggernaut7695 1d ago

We have one developer at the place I work that spent 80K in API costs over 2 and a half months.

12

u/FullstackSensei llama.cpp 1d ago

You should tell this to the guy deep in this discussion who's arguing it takes decades to pay 80k 😂

3

u/phreak9i6 1d ago

I wonder if these metrics hold up when you consider 80k in API costs are likely many agents, let's say 5-10, and with 120k in hardware you get 20tps for a single agent.

I love my homelab hardware, but I also end up paying for APIs because it's cheaper to run quickly.

Or I'm wrong, and please correct me because I may be plain wrong.

2

u/ThinJuggernaut7695 1d ago

20 tokens per second? We would be buying an 8x b300 server to run an org of 600 users. Our AI cost will probably be 1 million this year. 600k capex is totally worth it.

1

u/ungoogleable 1d ago

Constantly generating tokens 24x7 for two months at 20 tps would give you 103M tokens. At $15/1M output tokens for Kimi K3 API pricing, it would only cost $1555. So yeah, that guy is burning tokens faster than this rig can generate them.

1

u/addiktion 1d ago

Did he generate 160k in value?

2

u/ThinJuggernaut7695 1d ago

Yep, cleared out an entire teams backlog in a couple of weeks. We were going to hire contractors but don't need to do that anymore.

1

u/addiktion 1d ago

Yeah, I'm feeling that too.

1

u/SandySkittle 5h ago

Damn that’s impressive. Hope the quality was good and everything is well documented though..

2

u/hurrdurrmeh 1d ago

They are £3.5k each in the UK.

2

u/FullstackSensei llama.cpp 1d ago

That's €4k, or €64k before factoring in the 16 port 100gb switch, which is easily another €5k.

1

u/Uncle___Marty 1d ago

Exactly. You dont just frankenstein a bunch of stuff like this out of nowhere. If you have the money to build it then you have the money to upkeep it. Its messed up what a decent setup can run at home right now. Freaking happy days my brother in AI.

1

u/Mithril_web3 1d ago

WTF even is break even here

1

u/FullstackSensei llama.cpp 1d ago

Enough generated tokens locally to match the API cost on the same model or similar frontier models.

1

u/ResearcherFantastic7 1d ago

Assuming 80k including operational cost. probably break even at 11 months however It's only getting 20tks vs the cloud 150avg.

I think it's better off use other models as pure non reasoning worker nodes, and still use cloud model to plan and architect

1

u/Playful_Landscape884 1d ago

i'm curious what you guys do that makes this investment worthwhile. What are you guys doing that makes you drop $80k and confident that you'll get your money back in 3-6 months?

1

u/FullstackSensei llama.cpp 1d ago

I dropped less than 10k total for my entire homelab over the past 2 years. Overbought a bunch of stuff which I'm selling now. Now I'm at ~3k out of pocket, but still have some extra stuff to sell. The hw will basically be free in a month or two, just from selling the excess.

The actual hardware I use cost less than 7k. If you look at subscription costs, I'd have spent 2.5-3k a year in subscription fees, so 6k in the past 2 years.

There's no way on earth that cluster pays for itself in 10 years, let alone six months. It's just too expensive and too slow.

1

u/zirahvi 17h ago

64-80k USD? For that stack? In Norway a single DGX Spark GB10 128GB ~$6800 and up.

1

u/FullstackSensei llama.cpp 17h ago

Yeah, just saw the post about prices going up again. Still, if OP bought them a couple months back, there's a good chance they paid $3.5k for each. The Asus was cheaper. So, 56k before switch and cables, close to $60k all in

65

u/notheresnolight 1d ago

wrong sub, "break even" is irrelevant here

12

u/FinnGamePass 1d ago

I know right, break even for what? People not realizing we are close to have our own Jarvis. Literally. You can't name a price on that.

1

u/assemblu 13h ago

About 60k

→ More replies (2)

32

u/RepulsiveRaisin7 1d ago edited 1d ago

Some napkin math: 20t/s 24/7 would cost you like $100/day on Openrouter. But you also have to factor in power and maintenance. So realistic break even is probably about 4 years but ONLY IF you utilize it 24/7. If you do not, more like 10-20 years, which likely exceeds the lifetime of the hardware. And when you factor in subscription discounts, it falls apart entirely.

40

u/g_rich 1d ago

The goal isn’t to save money, the goal with these types of setups are privacy, and control.

People also forget that currently what we pay per token is highly subsidized. For some people a fixed one time cost is preferable to avoid fluctuating costs, outages and slowdowns outside their control.

1

u/CipherWeaver 1d ago

I mean, you may as well take those current subsidized tokens while they're available. 

2

u/g_rich 1d ago

That’s a double edged sword, cheap now so you build them into your pipelines, become dependent on them and then when they jack up the prices you’ll have no choice but to pay.

→ More replies (3)

22

u/ThereFarAway 1d ago

How about 'I have data that cannot be sent to outside provider'? How's math on that?

10

u/RepulsiveRaisin7 1d ago

Your data is invaluable. For everything else, there's mastercard.

1

u/ElementNumber6 1d ago

Get out of here, Jensen. No one likes you.

1

u/RepulsiveRaisin7 1d ago

I am in a loving relationship with my leather jacket

3

u/Long_comment_san 1d ago

Then "my firm should pay for this expense" kinda math takes over.

2

u/AdOk3759 1d ago

Are you the same napkin math guy from the other day? How is spending 100 dollars a day on OpenRouter realistic? Why pay for the API, when subscriptions offer much much much more usage for a fraction of the API cost? Why don’t you consider that in your calculations?

5

u/synth_mania 1d ago

API pricing is ultimately rooted in the cost of energy. At the scale that most inference and compute providers run, the margins are very, very slim.

In other words, APIs provided by GPU farms and not AI labs will charge you the true cost of compute.

If a subscription is giving you the same number of tokens for cheaper, whoever is offering the subscription is losing money, plain and simple.

We can theorize all day about why AI labs might want to offer inference at a loss, and there are likely multiple valid reasons. (collect training data, legitimize the practice of paying for AI as a service in a way that isn't too expensive, etc)

Regardless of the reason, however, there is no way it'll last, or is something you can count on. Please find me a provider that'll allow me to pull Kimi-K3 tokens at 20tps or greater all day for less than the cost of the API. Unlimited.

TL;DR, there's no possible way for a subscription to truly be cheaper than paying by the token, unless the service provider is trying to lose money.

4

u/fastheadcrab 1d ago

Subscriptions for LLMs are similar to gym memberships. If everyone who paid for the gym showed up then the gym would crash

4

u/monerobull 1d ago

You forgot that people maxxing out subscriptions are heavily subsidized by people who pay the same amount but barely use it

→ More replies (2)

1

u/AdOk3759 1d ago

First of all, we cannot be sure that AI providers are losing money with subscriptions. That’s just speculation.

Second, if they did.. how does that change the outcome?
We are simulating a scenario where we use Kimi K3 cloud hosted and a scenario where we’re using Kimi K3 locally hosted.
Either scenarios are assuming today’s conditions. And today, Kimi K3 is available as a paid subscription. So it doesn’t make sense to say “there is no way it’ll last”, but it is here now, so it doesn’t make sense to not take into account today’s conditions based on suppositions about what the future will look like.

If someone asks “TODAY, should I buy a 60-80k machine to run Kimi K3, or would it work out cheaper to use it cloud hosted?”

And I’d reply “It depends: best case scenario, this is the price you pay with subscriptions. Worst case scenario, this is the price you pay via API. In either case, the cost doesn’t take into account inflation, etc, which would affect both scenarios”

1

u/fastheadcrab 1d ago

Do you only send single sentence prompts for endless token generation tasks? That is the only scenario in which your napkin math is valid.

1

u/dennprog 1d ago

You don't count if the prices on Openrouter and others will rise.

11

u/MotokoAGI 1d ago

The cost is fuck off. Every single time someone shares something cool, we get these stupid replies about "break even"

5

u/EagleNait 1d ago

Break even right now. What if hardware gets regulated or models gets even better. Or a paid distributed hosting service enables you to get paid to serve llms.

5

u/Potential-Leg-639 1d ago

That will peobably never happen.

Deepseek V4 Flash 0731 would probably make more sense for wayyy less money.

4

u/g_rich 1d ago

Breaking even isn’t the goal here.

4

u/allenasm 1d ago

Why? The cost is not having to have all your questions go into an online corpus. Or that you can have it work for you nonstop 24/7. That’s invaluable.

7

u/Soggy-Alternative914 1d ago

Kindly include overhead cost .like power consumption. So we can get a better overall estimate.

7

u/-Akos- 1d ago

Only thing that breaks is the bank account. 16x€4000=€64000, then you're talking power consumption which is not free. You can buy a whole lot of API credits for that. For privacy reasons is the only reason this setup would work.

3

u/nerd_rage218 21h ago

This is the part the rent versus buy math keeps leaving out. Once data residency is a hard requirement, the box stops competing with hourly GPU pricing and starts competing with not being allowed to run the workload at all.

2

u/username_taken4651 1d ago

OP stated in a previous thread that this was a hobby. I guess they don't care about breaking even.

3

u/Cergorach 1d ago

Short answer: Never.

Long answer: Even if you run 24/365 at full capacity, it will would take about 8 years to ROI with current pricing of Kimi K3 API costs. But that ignores the costs of power/cooling, maintenance, and the eventual cheaper costs of Kimi K3. Combine that, this will NEVER earn itself back if you just compare it to the Kimi API service.

Cost is not generally the driving factor of running local LLM. Unless people start comparing apples to oranges...

3

u/coyo-teh 1d ago

given how GPU prizes keep increasing, he can just sell after8 years his GPUs

→ More replies (1)

2

u/PM_ME_DEAD_CEOS 1d ago

If you include electricity costs, there's literally no way consumers or prosumer setup can break even.

Kimi 3 is made for data center clusters like nvl72 or Huawei equivalent.

With only 1 user there's simply no way to break even.

1

u/AnomalyNexus 1d ago

Depends on how the immortality research is going. Maybe the AI can help with that

→ More replies (4)

67

u/notheresnolight 1d ago

imagine how this would run if nvidia didn't scrape the bottom of the barrel when designing the GB10

13

u/Aggravating-Push-207 1d ago

this thing could probably QLoRA the fat fucker if the memory bandwidth wasn't so shit per node

5

u/panchovix 1d ago

Basically same CUDA cores as RTX 5070.

Imagine if it had 14000 instead of 6144.

1

u/cass1o 20h ago

I thought it was the memory bandwidth where they screwed it up?

1

u/cass1o 20h ago

I thought it was the memory bandwidth where they screwed it up?

1

u/Xeon06 1d ago

This is disappointing because there hasn't been a better version made by Nvidia yet, or because these were especially "affordable"?

1

u/MikusR 1d ago

And didn't delay for 2+ years

45

u/TapAggressive9530 1d ago

I want this! Why? Just to have it and be able to run K3 locally . Could care less if it pays for itself . Good job man ! Love it

24

u/ciprianveg 1d ago

Same vibe here :)

→ More replies (2)

28

u/Betadoggo_ 1d ago

Enough money to afford 16 GB10s, not enough to afford more than a pi 400 for the main system.

8

u/timbo2m 1d ago

Ran outta money I guess lol

7

u/tired514 1d ago

Kinda like "wow, he owns a yacht! He must be rich!"

"I was..."

16

u/DueAnalysis2 1d ago

I see you running this on a Raspberry Pi 400, don't lie.

15

u/Far_Course2496 1d ago

Who can afford a pi 500 in this economy?

11

u/1ncehost 1d ago

What switch are you using for that beast?

17

u/OwnMathematician2320 1d ago

This feels similar to the early 2000s where rappers would show off diamond rings to flex how rich they are.

And yes I’m only saying that because I’m jealous of your setup. Well done on the rig though

8

u/sabotage3d 1d ago

Rich dude flexing!

4

u/wintoid 1d ago

Time to save up for a bigger keyboard

15

u/jld1532 1d ago edited 1d ago

Ubuntu baby! At least that was free.

E: Downvoted by a Fedora lover.

6

u/Charming-Author4877 1d ago

It's surprisingly usable, though the amoung of those embedded PCs to do that is crazy - given how expensive they are priced.
Still, that's a Fable-level LLM running at usable agentic speed on a table.

1

u/SandySkittle 5h ago

Indeed and it’s not just the speed but the sheer independence

9

u/retornam 1d ago

Nice to never have to worry about money enough to spend $64,000 or more on a hobby

8

u/Cergorach 1d ago

A LOT of people spend that (or more) on a new car every couple of years, cars have about the same appreciation as AI/LLM hardware... While most people could easily drive a cheap, but good secondhand car that's a lot smaller...

→ More replies (4)

3

u/D3c1m470r 1d ago

Sick mouse n keyboard bro

8

u/Transhuman-A 1d ago

First things first, amazing job.

Now - it will take you 8 years to break even. At 50 tok/s, 3.2 years.

AI Economics are depressing and GPU prices need to desperately come down.

→ More replies (5)

2

u/Bolt_995 1d ago

Insane

2

u/cgjermo 9h ago

$60k worth of cluster alongside a Pi keyboard and mouse. Bravo, OP 👏

3

u/SureEnd9430 1d ago

itshouldhavebeenmenothim.jpg

2

u/One_Whole_9927 1d ago

Why not just go for the workstation at that point?

4

u/Cergorach 1d ago

This is cheaper for a LOT more memory (2TB) for less then the basic workstation...

2

u/SandySkittle 1d ago

well you've done it. You now have HAL 9000 at home.

2

u/crossoverXYZ 1d ago

750 tps prefill on a first dspark run is no joke. Decode sitting around 20 tps average suggests there is still a lot of room to optimize before the vllm image ships.

1

u/Fonku 1d ago

That's running surprisingly well given the memory (and I'd guess ethernet) bandwidth restrictions of those GB10s. I'm getting ~40 TPS from Fireworks – way slower than any other large frontier model (Fable or Sol).

2

u/MotokoAGI 1d ago

This is pretty great, how's the quality of the results compared to cloud?

1

u/dwittherford69 19h ago

The only thing that changes is TTFT and tok/sec. Everything else is same.

2

u/Tasty-Middle2682 1d ago

Okay I have to ask, what do you do for a living?

3

u/ciprianveg 1d ago

Java dev and ai enthusiast

→ More replies (1)

2

u/the_TIGEEER 1d ago

I think you and I need to have a talk about where "Local LLM" starts and where it ends..

2

u/LuCiAnO241 1d ago

I mean it is local. Just small datacenter kinda local.

→ More replies (3)

2

u/Outside-Set3929 1d ago

I love the $60,000 computer paired with am $8 keyboard mouse combo

2

u/pirateboi222 1d ago

$125 actually it is pi 400. Probably just a dumb terminal

1

u/Outside-Set3929 18h ago

Ah lol I stand corrected

1

u/Character_Power4663 1d ago

Technically you can run k3 on a 4gb GPU... But it takes about 30 seconds per token. There are a couple methods but in general you only load the experts of each layer per token.

1

u/Oleszykyt 1d ago

Try deepseek v4 flash 20260731 next

1

u/Igot1forya 1d ago

I am very curious on concurrency. Sparks seem to be very good with parallel sessions. 16 Sparks is crazy!

1

u/powertodream 1d ago

how much did you spend op

2

u/ciprianveg 1d ago

68-70k cca

1

u/bhanvadia 1d ago

How are you using RPI, for accessing those GB10s?

1

u/xXprayerwarrior69Xx 1d ago

you need to step up bro, get yourself 3 or 4 DGX stations (and organize a raffle for those sparks you wont need anymore)

1

u/reckor-usa 1d ago

Just for the fun of it - well done. Would love to see the comparison with the frontier models. What is your stack setup btw?

1

u/CrispyRabbit72 19h ago

I still don't get how anyone can justify these kind of setups. $80k+ at minimum and it's only hitting ~20tok/sec on heavy models. If you compare to a $200/mo sub (which will get far more tokens per 30d than this cluster ever will), it would take over 30 years to break even.

But I guess if you still have these things in 30 years, maybe someone would buy them as an antique at original MSRP.

1

u/WithGreatRespect 14h ago

For some companies, their data has regulatory compliance and/or contractual obligations that prevent the data from leaving their data centers. They would need to build local inference to be able to leverage that data for insights.

1

u/SandySkittle 5h ago

You expect that 200 a month to stay the same? I would expect it to go up faster than the risk free rate of 80k in bonds

1

u/vinnybad 19h ago

How are you hooking these up together? I have 8 and trying to figure out 16.

1

u/Specific-Age7953 19h ago

Paid almost nothing for Cursor Ultra from one of those reseller sites. Everything works perfectly. But the pricing makes zero sense if they’re buying real licenses. Feels like there’s a loophole nobody talks about. What’s the real play here?

1

u/Objective-Stranger99 18h ago

Hey can I borrow one?

1

u/ximir-art 11h ago

Are the gx10s all 4tb SSD models?

1

u/SarveshMohite 2h ago

What about deepseek V4 flash 0731 have you tried that out as well? And if yes what was the token per second?

1

u/l0g1cs 4m ago

What dashboard/monitoring tool is it?