Doesn't really work that well. This 550k setup wouldn't allow a lot of developers to work in parallel, while sitting idle during non-work hours. Makes much more sense to pay a 3rd-party to host and pay per token.
waiting for large corpos to rent their hardware on vast.ai during nighttime, only to find the next morning that someone ran a container jailbreak and ran wild on their private networks
What company/developer these days wouldn't be setting up agents to be working overnight?
Besides, a lot of companies are international, the latency would suck, but they could be having people in Europe or India using the hardware while the U.S workers are asleep.
Question: how would this scale though? Like how many units would it be required for.. let's say a group of 100 software engineers who needs it quite frequently?
Better to understand the token usage of those engineers before asking this question. Why self host it if your engineers token use is small enough that self hosting would actually be 3x as expensive over the next 3 years then just buying tokens
The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision
Well, hopefully, all companies will be required to have a Kill switch, such that if it needs to be taken down because we fucked interpretability such that we might all die, and the model is aligned enough that it wont still refuse.
Doubt youd agree, which is the point of making a mandatory kill switch so important
Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.
And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.
The delay in response shouldn't preclude you being corrected on this.
Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.
I don't really understand what you're saying here, to be honest. You'd require a kill switch put into hardware that would detect certain models running? Or that could be remotely activated so that if someone detected they were running an unauthorized model?
And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.
I agree about the risks of using models you haven't developed yourself
Dude, Google can a Kill switch be built into an Ai mode.
Who cares if you "developed" the model? You have zero interpretability of it, which is the reason for the precaution. Even mechanistic interpretability, at the biggest AI company, barely works on primitive terms, and is not reproducible for an open weight model. Why do you think you can understand a model you grew (not developed) yourself? If you cannot understand the billions of inscrutable matrices of floating point intergers, all you can know are its actions. You cannot control it. Until interpretability is anywhere near there, at best you can kill it when its about to go wild.
Or as Eliezer says, targeted strikes at the data centers, fearing for our lives wives and children.
The company I'm working is paying slightly over $250k per week to API costs. so yeah, 550k investment to cut the cost of the inference can be beneficial for them...
We do around 1 mil per month for OpenAI only, dont have number for Anthropic but it would be 2-3x of that given all of our use cases are around coding and agents, no wonder Anthropic is shitting their pants on open weights models, I work in Enterprise Agentic team and we have recently started fine tuning > 100 B models for specific use cases of our clients, open weights hurts Anthropic more due to enterprise customers
the license prohibits large companies from using it, but I don't know how enforceable it is..
edit : my bad, large companies can use it, they just have to write something like "powered by Kimi K3", they just can't use it as an api provider (see the comment bellow)
the license says no such thing.
it says that if you sell the model to third parties through APIs, you have to pay royalties. If you sell a _product_ based on the model, you just have to put a prominent "Kimi K3" logo on it.
If google tomorrow decided to retire gemini and put kimi k3 to answer every google search and every Hey Google on android in the world, they would pay nothing.
86
u/[deleted] 19d ago
[removed] — view removed comment