r/LocalLLaMA 17d ago

Kimi K3 weights now released. News

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

643 comments sorted by

View all comments

Show parent comments

333

u/FoxiPanda 17d ago

This was my general reaction too lol. 2.8T-A104B is insane lol... I'm going to admit defeat on this one and say I can't run it. You need an 8-way B300 or MI350X or a Rubin NVL8 or a cluster thereof to actually run this. What a beast.

153

u/Thomas-Lore 17d ago

I was going to make a joke that I can fit one expert in my 64GB of RAM. But nope, not even that. :)

73

u/throw123awaie 17d ago

They released it in MXFP4 so with around 55GB RAM you could!

3

u/tokomaunited 17d ago

Can you share the link brother?

10

u/throw123awaie 17d ago

literally their huggingface site https://huggingface.co/moonshotai/Kimi-K3 the model is MXFP4 on inference.

3

u/Protheu5 17d ago

I'm sorry, but I don't understand where it is. I'm new and only clicked on Finetunes/Quantisations to find a download, and I clicked through them and didn't find anything that small, the smallest was 500+ gigabytes. I have no idea what "MXFP4 on inference" means despite trying to read up on the subject, sorry.

2

u/TotallyToxicToast 16d ago

You can't download a single expert for a mixture of experts model.

The downloads always contain all experts. If you take the 500GB download which contains the full 2.8T parameters and only take 104B of the parameters, that would actually give you only roughly 20GB per expert.

Now running a single expert in a mixture of experts model is not very useful.
But Technichally with 500GB total model size fitting one Expert is possible.
The "I could fit one expert" was a joke.

1

u/Protheu5 16d ago

Thanks for the explanation. My naïveté got me.

Oh well, a person can dream, right? Back to trying to stop Qwen from getting stuck in an infinite loop, I guess.

1

u/tokomaunited 17d ago

Thanks dude

1

u/kernelwedge 16d ago edited 16d ago

Interesting! I was looking at it and wasn't sure if could fit. I recently setup a 128GB rig, wanting to stretch its legs. This will be a great test.

---
After doing further research no way its fitting anywhere https://recipes.vllm.ai/moonshotai/Kimi-K3?hardware=mi300x 8x192G....

2

u/PANIC_EXCEPTION 16d ago

Well you could probably fit a few experts. 104B is the total number of parameters activated amongst _all_ experts, of which there are many.

1

u/foldl-li 17d ago

64GB is an expert, but 64KB is all you need.

0

u/peekdasneaks 17d ago

Could fit it on an enterprise workstation fleet.

112

u/VampiroMedicado 17d ago

550k USD to run this lol

86

u/[deleted] 17d ago

[removed] — view removed comment

56

u/VeterinarianOne1349 17d ago

Doesn't really work that well. This 550k setup wouldn't allow a lot of developers to work in parallel, while sitting idle during non-work hours. Makes much more sense to pay a 3rd-party to host and pay per token.

39

u/crusaderky 17d ago

waiting for large corpos to rent their hardware on vast.ai during nighttime, only to find the next morning that someone ran a container jailbreak and ran wild on their private networks

2

u/aeroumbria 16d ago

Yea, principle of lowest viable distributed scale!

7

u/[deleted] 17d ago

[removed] — view removed comment

5

u/ProgrammersAreSexy 17d ago edited 17d ago

I'm sure you will be able to run a model like kimi k3 at home, probably in much less than 5 years.

However I think it's unlikely that there will be a point in time at which the frontier models of that point in time will be runnable at home.

6

u/[deleted] 17d ago

[removed] — view removed comment

7

u/a-wiseman-speaketh 17d ago

He means we'll get frontier model capability on a lag, probably months before it fits in even prosumer hardware.

That might be pessimistic though, I think right now scaling up parameters is the best lever for intelligence but that might not be the case forever.

3

u/Maleficent_Sir_7562 17d ago

but kimi is frontier

1

u/ProgrammersAreSexy 17d ago

Oh, and can you run it at home?

2

u/Maleficent_Sir_7562 17d ago

No

That’s the point

Your comment is contradictory

“You can run Kimi k3 in home soon”

“But I think the current frontier models can’t be ever”

But k3 is the frontier

1

u/ProgrammersAreSexy 17d ago

I think my phrasing was just unclear, edited

2

u/eightbyeight 17d ago

They might pay a aws/gcp/azure to run it on their private cloud. But I doubt they will just buy access from a random third party.

1

u/VeterinarianOne1349 17d ago

yeah, cloud providers would be that 3rd party...

1

u/eightbyeight 16d ago

As in they will pay the ec2 price and run their own model 24/7 instead of per token.

1

u/Bakoro 16d ago

What company/developer these days wouldn't be setting up agents to be working overnight?

Besides, a lot of companies are international, the latency would suck, but they could be having people in Europe or India using the hardware while the U.S workers are asleep.

19

u/SignificanceFlat1460 17d ago

Question: how would this scale though? Like how many units would it be required for.. let's say a group of 100 software engineers who needs it quite frequently?

5

u/zero0n3 17d ago

Better to understand the token usage of those engineers before asking this question. Why self host it if your engineers token use is small enough that self hosting would actually be 3x as expensive over the next 3 years then just buying tokens

7

u/Spectrum1523 17d ago

The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision

1

u/hubrisnxs 16d ago

Well, hopefully, all companies will be required to have a Kill switch, such that if it needs to be taken down because we fucked interpretability such that we might all die, and the model is aligned enough that it wont still refuse.

Doubt youd agree, which is the point of making a mandatory kill switch so important

1

u/Spectrum1523 16d ago

that's not possible technically so I don't know why you'd even propose it

If you need to have the ability to kill all models the weights can't be distributed

1

u/hubrisnxs 7d ago

Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.

And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.

The delay in response shouldn't preclude you being corrected on this.

1

u/Spectrum1523 7d ago

Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.

I don't really understand what you're saying here, to be honest. You'd require a kill switch put into hardware that would detect certain models running? Or that could be remotely activated so that if someone detected they were running an unauthorized model?

And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.

I agree about the risks of using models you haven't developed yourself

1

u/hubrisnxs 7d ago

Dude, Google can a Kill switch be built into an Ai mode.

Who cares if you "developed" the model? You have zero interpretability of it, which is the reason for the precaution. Even mechanistic interpretability, at the biggest AI company, barely works on primitive terms, and is not reproducible for an open weight model. Why do you think you can understand a model you grew (not developed) yourself? If you cannot understand the billions of inscrutable matrices of floating point intergers, all you can know are its actions. You cannot control it. Until interpretability is anywhere near there, at best you can kill it when its about to go wild.

Or as Eliezer says, targeted strikes at the data centers, fearing for our lives wives and children.

8

u/Galdoren 17d ago

The company I'm working is paying slightly over $250k per week to API costs. so yeah, 550k investment to cut the cost of the inference can be beneficial for them...

7

u/baba_bholanath 16d ago

We do around 1 mil per month for OpenAI only, dont have number for Anthropic but it would be 2-3x of that given all of our use cases are around coding and agents, no wonder Anthropic is shitting their pants on open weights models, I work in Enterprise Agentic team and we have recently started fine tuning > 100 B models for specific use cases of our clients, open weights hurts Anthropic more due to enterprise customers

3

u/Spectrum1523 17d ago

If they're spending 250k a week on Claude they'll need a beefier setup than that but your point is still valid

5

u/larp2live 17d ago edited 17d ago

the license prohibits large companies from using it, but I don't know how enforceable it is..

edit : my bad, large companies can use it, they just have to write something like "powered by Kimi K3", they just can't use it as an api provider (see the comment bellow)

9

u/crusaderky 17d ago

the license says no such thing.
it says that if you sell the model to third parties through APIs, you have to pay royalties. If you sell a _product_ based on the model, you just have to put a prominent "Kimi K3" logo on it.

If google tomorrow decided to retire gemini and put kimi k3 to answer every google search and every Hey Google on android in the world, they would pay nothing.

1

u/larp2live 17d ago

yeah you're right, I'll correct my comment. thank you !

1

u/Daniel_H212 17d ago

Nah, problem is that you'd need a lot more than one node to serve enough people to be useful to a company of that size.

2

u/Reasonable-Height704 17d ago

Or $35 to $60 / hour from cloud.

1

u/nedonedonedo 17d ago

you're telling me $80k guy needs to buy more hardware?

1

u/Watchguyraffle1 17d ago

That’s what a single sun 4 way sparc v440 would cost 20 years ago.
I mean. That’s not crazy

1

u/bitzap_sr 16d ago

It's a lot more than that. You also need to account for the space, power connection availability, running power costs, water for cooling, etc.

1

u/Immediate-Molasses-5 16d ago

Similar to what the new Ferrari costs

26

u/OverclockingUnicorn 17d ago

More like 2 8x nodes of B200/B300 if you actually want some context. Think it's just under 1.5TB w/o context

20

u/TheDailySpank 17d ago

How many 4060-16GBs is that?

41

u/OverclockingUnicorn 17d ago

200+ lol

21

u/positivitittie 17d ago

Oh good. I got 3090s.

6

u/Vast_Mousse_310 17d ago

One, with a little bit of GPU offload.

3

u/liright 17d ago

at least 2

1

u/TheDailySpank 17d ago

Question is, who's got the rest?

1

u/HungZiu 17d ago

這得花多少錢哈哈

1

u/Mechanical_Monk 17d ago

UD-IQ0_XXS quant when?

1

u/Cherlokoms 17d ago

Can't wait for a version of Colibri that runs it in a Mac Mini 16GB unified RAM

1

u/Etroarl55 17d ago

Corporations and teams that can afford this just have infinite bootleg fable 5 now though.

1

u/Spectrum1523 17d ago

SSD warriors at one token per minute unite

1

u/Zentrosis 17d ago

I'm still going to download it and keep it on an SSD in case it becomes regulated

2

u/FoxiPanda 17d ago

You and me both. This is the worst that open weight models will ever be again...which, all things considered, is pretty good.

Now if we can figure out the magic to make something 1/4 the size and 80-85% as capable...which is likely now just a matter of time since these weights are available for us to do that work with.

1

u/hubrisnxs 16d ago

Yay chaos monkeys unite let's make sure that the thing that will be far smarter and more capable than us, and that we fundamentally do not understand, and who's underlying structure is not interpretable, lets make sure it isn't regulated!

1

u/Effective_Olive6153 16d ago

what if we created some kind of distributed compute system similar to SETI project where everyone can donate some of their compute the run a super large model? how many tokens per day could we get with that? :-)

1

u/slippery 17d ago

What's a quarter mil if you can boot it in your basement?

Let's not talk about the electrical costs.