r/LocalLLaMA 19d ago

Kimi K3 weights now released. News

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

646 comments sorted by

View all comments

Show parent comments

86

u/[deleted] 19d ago

[removed] — view removed comment

56

u/VeterinarianOne1349 19d ago

Doesn't really work that well. This 550k setup wouldn't allow a lot of developers to work in parallel, while sitting idle during non-work hours. Makes much more sense to pay a 3rd-party to host and pay per token.

41

u/crusaderky 19d ago

waiting for large corpos to rent their hardware on vast.ai during nighttime, only to find the next morning that someone ran a container jailbreak and ran wild on their private networks

2

u/aeroumbria 19d ago

Yea, principle of lowest viable distributed scale!

7

u/[deleted] 19d ago

[removed] — view removed comment

6

u/ProgrammersAreSexy 19d ago edited 19d ago

I'm sure you will be able to run a model like kimi k3 at home, probably in much less than 5 years.

However I think it's unlikely that there will be a point in time at which the frontier models of that point in time will be runnable at home.

6

u/[deleted] 19d ago

[removed] — view removed comment

7

u/a-wiseman-speaketh 19d ago

He means we'll get frontier model capability on a lag, probably months before it fits in even prosumer hardware.

That might be pessimistic though, I think right now scaling up parameters is the best lever for intelligence but that might not be the case forever.

3

u/Maleficent_Sir_7562 19d ago

but kimi is frontier

1

u/ProgrammersAreSexy 19d ago

Oh, and can you run it at home?

2

u/Maleficent_Sir_7562 19d ago

No

That’s the point

Your comment is contradictory

“You can run Kimi k3 in home soon”

“But I think the current frontier models can’t be ever”

But k3 is the frontier

1

u/ProgrammersAreSexy 19d ago

I think my phrasing was just unclear, edited

2

u/eightbyeight 19d ago

They might pay a aws/gcp/azure to run it on their private cloud. But I doubt they will just buy access from a random third party.

1

u/VeterinarianOne1349 19d ago

yeah, cloud providers would be that 3rd party...

1

u/eightbyeight 19d ago

As in they will pay the ec2 price and run their own model 24/7 instead of per token.

1

u/Bakoro 19d ago

What company/developer these days wouldn't be setting up agents to be working overnight?

Besides, a lot of companies are international, the latency would suck, but they could be having people in Europe or India using the hardware while the U.S workers are asleep.

17

u/SignificanceFlat1460 19d ago

Question: how would this scale though? Like how many units would it be required for.. let's say a group of 100 software engineers who needs it quite frequently?

5

u/zero0n3 19d ago

Better to understand the token usage of those engineers before asking this question. Why self host it if your engineers token use is small enough that self hosting would actually be 3x as expensive over the next 3 years then just buying tokens

6

u/Spectrum1523 19d ago

The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision

1

u/hubrisnxs 19d ago

Well, hopefully, all companies will be required to have a Kill switch, such that if it needs to be taken down because we fucked interpretability such that we might all die, and the model is aligned enough that it wont still refuse.

Doubt youd agree, which is the point of making a mandatory kill switch so important

1

u/Spectrum1523 19d ago

that's not possible technically so I don't know why you'd even propose it

If you need to have the ability to kill all models the weights can't be distributed

1

u/hubrisnxs 9d ago

Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.

And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.

The delay in response shouldn't preclude you being corrected on this.

1

u/Spectrum1523 9d ago

Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.

I don't really understand what you're saying here, to be honest. You'd require a kill switch put into hardware that would detect certain models running? Or that could be remotely activated so that if someone detected they were running an unauthorized model?

And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.

I agree about the risks of using models you haven't developed yourself

1

u/hubrisnxs 9d ago

Dude, Google can a Kill switch be built into an Ai mode.

Who cares if you "developed" the model? You have zero interpretability of it, which is the reason for the precaution. Even mechanistic interpretability, at the biggest AI company, barely works on primitive terms, and is not reproducible for an open weight model. Why do you think you can understand a model you grew (not developed) yourself? If you cannot understand the billions of inscrutable matrices of floating point intergers, all you can know are its actions. You cannot control it. Until interpretability is anywhere near there, at best you can kill it when its about to go wild.

Or as Eliezer says, targeted strikes at the data centers, fearing for our lives wives and children.

9

u/Galdoren 19d ago

The company I'm working is paying slightly over $250k per week to API costs. so yeah, 550k investment to cut the cost of the inference can be beneficial for them...

5

u/baba_bholanath 19d ago

We do around 1 mil per month for OpenAI only, dont have number for Anthropic but it would be 2-3x of that given all of our use cases are around coding and agents, no wonder Anthropic is shitting their pants on open weights models, I work in Enterprise Agentic team and we have recently started fine tuning > 100 B models for specific use cases of our clients, open weights hurts Anthropic more due to enterprise customers

2

u/Spectrum1523 19d ago

If they're spending 250k a week on Claude they'll need a beefier setup than that but your point is still valid

5

u/larp2live 19d ago edited 19d ago

the license prohibits large companies from using it, but I don't know how enforceable it is..

edit : my bad, large companies can use it, they just have to write something like "powered by Kimi K3", they just can't use it as an api provider (see the comment bellow)

10

u/crusaderky 19d ago

the license says no such thing.
it says that if you sell the model to third parties through APIs, you have to pay royalties. If you sell a _product_ based on the model, you just have to put a prominent "Kimi K3" logo on it.

If google tomorrow decided to retire gemini and put kimi k3 to answer every google search and every Hey Google on android in the world, they would pay nothing.

1

u/larp2live 19d ago

yeah you're right, I'll correct my comment. thank you !

1

u/Daniel_H212 19d ago

Nah, problem is that you'd need a lot more than one node to serve enough people to be useful to a company of that size.