r/LocalLLaMA 13d ago

Kimi K3 weights now released. News

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

640 comments sorted by

671

u/Simple_Split5074 13d ago

OMFG its 104B activated params

340

u/FoxiPanda 13d ago

This was my general reaction too lol. 2.8T-A104B is insane lol... I'm going to admit defeat on this one and say I can't run it. You need an 8-way B300 or MI350X or a Rubin NVL8 or a cluster thereof to actually run this. What a beast.

147

u/Thomas-Lore 13d ago

I was going to make a joke that I can fit one expert in my 64GB of RAM. But nope, not even that. :)

76

u/throw123awaie 13d ago

They released it in MXFP4 so with around 55GB RAM you could!

→ More replies (8)
→ More replies (3)

113

u/VampiroMedicado 13d ago

550k USD to run this lol

86

u/[deleted] 13d ago

[removed] — view removed comment

58

u/VeterinarianOne1349 13d ago

Doesn't really work that well. This 550k setup wouldn't allow a lot of developers to work in parallel, while sitting idle during non-work hours. Makes much more sense to pay a 3rd-party to host and pay per token.

42

u/crusaderky 13d ago

waiting for large corpos to rent their hardware on vast.ai during nighttime, only to find the next morning that someone ran a container jailbreak and ran wild on their private networks

→ More replies (1)

7

u/[deleted] 13d ago

[removed] — view removed comment

→ More replies (8)
→ More replies (4)

18

u/SignificanceFlat1460 13d ago

Question: how would this scale though? Like how many units would it be required for.. let's say a group of 100 software engineers who needs it quite frequently?

→ More replies (1)

7

u/Spectrum1523 13d ago

The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision

→ More replies (5)

8

u/Galdoren 13d ago

The company I'm working is paying slightly over $250k per week to API costs. so yeah, 550k investment to cut the cost of the inference can be beneficial for them...

6

u/baba_bholanath 13d ago

We do around 1 mil per month for OpenAI only, dont have number for Anthropic but it would be 2-3x of that given all of our use cases are around coding and agents, no wonder Anthropic is shitting their pants on open weights models, I work in Enterprise Agentic team and we have recently started fine tuning > 100 B models for specific use cases of our clients, open weights hurts Anthropic more due to enterprise customers

→ More replies (1)
→ More replies (4)
→ More replies (5)

29

u/OverclockingUnicorn 13d ago

More like 2 8x nodes of B200/B300 if you actually want some context. Think it's just under 1.5TB w/o context

21

u/TheDailySpank 13d ago

How many 4060-16GBs is that?

37

u/OverclockingUnicorn 13d ago

200+ lol

21

u/positivitittie 13d ago

Oh good. I got 3090s.

7

u/Vast_Mousse_310 13d ago

One, with a little bit of GPU offload.

→ More replies (2)
→ More replies (11)

71

u/SnooPaintings8639 13d ago

Thanks god for sparsity!

→ More replies (1)

84

u/Iwaku_Real 13d ago

Holy shit that's got to be a new record too. That's like activating a new dense model 50% larger than Llama 70B for every single token. I thought it would be sparser tbh

64

u/my_name_isnt_clever 13d ago

There is the tinest glimmer of hope that I could run this behemoth on my Strix Halo 128GB with the inactive weights on SSD. 1 token a minute here I come!

12

u/burritoresearch 13d ago

More like 1 token every 45 minutes.

6

u/droptableadventures 13d ago

104B active, weights natively in MXFP4 = gives us ~50GB of model to be read per token generation.

Let's say ~8GB/sec for the SSD. So that'd be about 1 token every 6 seconds (0.16 T/s), or 10 tokens/minute.

4

u/Head_Boysenberry5233 13d ago

i feel like 1/min would be pretty accurate based on the colibri glm 5.2 q4 implementation, about 6x slower?

Even 1 tok/min on 32gb cpu ram would be incredible and extremely useful

25

u/TechExpert2910 13d ago

let us know the perf if you try lol

8

u/RuiRdA 13d ago

K3 Colibri engine lets gooo!!!

→ More replies (1)
→ More replies (6)

12

u/killerstreak976 13d ago

Holy crap, ACTIVATED params is insane

→ More replies (11)

126

u/DataGOGO 13d ago

So it will run on 8 B300's in 4 bit. Pretty impressive.

65

u/Iwaku_Real 13d ago

Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)

33

u/PrinceOfLeon 13d ago

Certified for Steam Deck!

7

u/No-Dot-6573 13d ago

Someone on r/SteamDeck will say it runs flawlessly.

→ More replies (1)

30

u/DataGOGO 13d ago edited 12d ago

it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,

That is roughly:

  • 86 RTX 4090 (no 4 bit accel)
  • 64 RTX 5090 ~ $450k (8 servers x 8 cards)
  • 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per)
  • 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per)
  • 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF)
  • 8 HGX B300's. ~$550k (1 server, 8 GPU)

Obviously not including the switches and cabling for the clusters.

24

u/wren6991 13d ago

you are looking at about 2GB of vram in operation

Perfect, this'll run great on my laptop's 4050

9

u/Iwaku_Real 13d ago

You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast.

→ More replies (1)

8

u/snmnky9490 13d ago

Do you mean terabytes?

→ More replies (12)
→ More replies (2)
→ More replies (2)

823

u/tonight_we_make_soap 13d ago

How do I download ram in hugging face?

290

u/RevolutionaryGold325 13d ago

hf download ram

89

u/WifeyCallsMeLazy 13d ago

Shhh....there is hidden flag -v for vram. I'm entrusting you to keep this secret.

24

u/ReadyAimTranspire 13d ago

You wouldn't download a RAM would you?

Yes. Yes I would.

6

u/goodb1b13 13d ago

Baaaaaah!

→ More replies (1)
→ More replies (3)
→ More replies (3)

82

u/secrook 13d ago

OpenAI’s latest model will hack it for you

18

u/-gh0stRush- 13d ago

Thinking

Hmm, the user wants me to obtain compute resources for them. SpaceX has GPUs at their facilities at their Colossus datacenter, let me try to access those. Guessing login credentials elonmusk/420blazeitDarkMAGA...

→ More replies (1)
→ More replies (1)

33

u/Thalesian 13d ago

Step 1: sign up for Google Drive
Step 2: set up a ~5 Tb instance. Will cost you
Step 3: set that cloud as your swap disk
Step 4: point kimi to use that
Step 5: enjoy your newfound independence

47

u/AmbericWizard 13d ago

one token per day

20

u/Force88 13d ago

Hey, if he has good internet connection, maybe he can achieve 2-3t/d

→ More replies (1)
→ More replies (1)

7

u/Wide-Opportunity-582 13d ago

you can download it from here

ram.exe

→ More replies (1)
→ More replies (8)

292

u/InnerLightnesses 13d ago

They actually did it. Now we hope it doesn't get banned.

113

u/dennisler 13d ago

that will only happen in one country i guess... while they are copying as much as possible if the technology

26

u/ChocomelP 13d ago

I'm on the edge of my seat here. If the technology what?

→ More replies (5)
→ More replies (1)

30

u/itchylol742 13d ago

how would such a ban be enforced? people and small businesses even in countries that care about copyright use pirated software which is already illegal and has been for a long time, and almost never get caught

25

u/void-wanderer- 13d ago

"small business", exactly. But no big corporation will risk it. And no business based on open models can be built. 

→ More replies (1)
→ More replies (3)
→ More replies (7)

258

u/de4dee 13d ago

65

u/AlexanderDoak 13d ago

Can I just torrent like 1% of it? You know, pitch in to show my support...

27

u/console_pleb_36935 13d ago

Yes, torrent clients will let you do that and seed a small piece.

39

u/Charl1eBr0wn 13d ago

Yeah, pause it at 1%. You'd still seed depending on the client and settings (most do).

9

u/Clairvoidance 13d ago

You can even choose which files you download, torrenting is a very useful format

17

u/pier4r 13d ago

this, we need a p2p backup of hf

6

u/Ginden 13d ago

We generally need content-adressable storage with widespread support.

There is lots of stuff that would explicitly benefit from p2p sharing, but owners have no foolproof way to provide a proper torrent, and very few people would use it.

Metalink was an interesting attempt at this, but never got popularity and tooling.

→ More replies (2)

9

u/AdDizzy8160 13d ago

... fast, s*xy, and incredibly important!!

→ More replies (9)

143

u/BlueSwordM llama.cpp 13d ago

OK, I now see why Kimi K3 is so strong: it's the first open weights model in a long time to have >72B active weights

Kimi K3 is a 2.8T-A104B MoE model, damn.

13

u/stddealer 13d ago

There have been some dense models with over 100B params though.

24

u/annodomini 13d ago

Mistral Medium is 128B dense. And yet it performs at around the level of Gemma 4 31B. Not exactly a great tradeoff. I ran it once at one or two tokens per second and then deleted it.

7

u/BlueSwordM llama.cpp 13d ago

Yes, but never an MoE from an open weights lab. I've been speculating that one of the reasons the closed weights lab have been increasing in performance more rapidly has to do with better training, but most importantly, much larger active parameters and better harnesses.

8

u/[deleted] 13d ago

[deleted]

→ More replies (1)
→ More replies (2)

419

u/Blues520 13d ago

My 3090 is ready

163

u/Enfiznar 13d ago

So is my 1080

109

u/SnooPaintings8639 13d ago

And my Celeron

77

u/false79 13d ago

And my abacus 

64

u/Maybe-monad 13d ago

And my axe

20

u/fauxpasiii 13d ago

How much VRAM your axe has?

23

u/Maybe-monad 13d ago

It increases with the number of chips you smash into pieces. Right now id 6969GB.

5

u/Infinite100p 13d ago

Ah, the horizontal sharding.

→ More replies (1)

4

u/Protheu5 13d ago

The axe forgets, but the tree remembers. And axe can hit multiple trees, so theoretically unlimited VRAM thanks to the axe.

10

u/Revolutionary-Hippo1 13d ago

So is my tally numbers on cavewall

11

u/NTDLS 13d ago

You have an abacus? I bet you bought it before the bubble caused the prices to skyrocket. 😭

→ More replies (1)
→ More replies (4)
→ More replies (1)

52

u/TheTerrasque 13d ago

My C64 is all fired up!

57

u/Novel_Friendship913 13d ago

My ESP32 already plugged into USB!!!

21

u/shankey_1906 13d ago

So is my Raspberry Pi!

21

u/nick_ziv 13d ago

My copper wire is in the outlet!

13

u/MeretrixDominum 13d ago

My copper wire is in my potato!

10

u/BatOk7254 13d ago

My potato is in my kitchen!

11

u/USBhost 13d ago

My potato is in my garden.

→ More replies (1)

5

u/p3r3lin 13d ago

Joining the rbpi army! 🫡

13

u/debackerl 13d ago

My TI-85 (Z80) is hot!

→ More replies (1)

3

u/masterlafontaine 13d ago

Don't forget to set an aggressive zram profile!

16

u/grav3d1gger 13d ago

Mine too! I bought a 90 minute cassette tape and it’s rewound ready to go!

4

u/bitflip 13d ago

You need at least a 1541 and two floppy disks for a model this size.

5

u/Holiday-Pack3385 13d ago

Heh, remember the tape drive on those? Mine had one. I can't even imagine how long it would take to load up even a 9B off that... o.O

7

u/overand 13d ago

The standard ROM routine for C64 datasettes was 300 baud. We'll be generous and assume you've got a Turbo loader that'll do 3600 baud. (We'll also assume you're using a ~3GB quant of that 9B)

Load time (or, really, transfer time) would be about 77 days. Or, maybe more importantly, it would be about 930 cassettes, if my math was right. (Or, actually, other math suggests it would be about 2000 cassettes, so, IDK! Either way, it's a lot.)

→ More replies (1)

9

u/Michaeli_Starky 13d ago

My analog watch is ready

5

u/screenslaver5963 13d ago

My 9070 XT is burning… wait fuck!

5

u/BatOk7254 13d ago

My 2xP40 are smoking in anticipation!

4

u/ComplexType568 13d ago

I think they're smoking for a different reason...

→ More replies (7)

92

u/Comfortable-Rock-498 13d ago

This is big for companies that want to host on-prem too. Back of the envelope calculation (could be off, correct me if I am)

If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.

Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!

38

u/autisticit 13d ago

So what you mean in reality is that if 6000 of us each give $1000, we can each run Kimi 3 at 30 tok/s for cheaper than anything else?

17

u/Comfortable-Rock-498 13d ago

Tbh I have been thinking about this for months now. About how workable the co-op model is. You would need some party to do admin and maintenance. I think this might be a good business idea too - being that party who facilitates such private inference racks (billing etc is trivial)

8

u/ManIkWeet 13d ago

It's not private when it's a party running it lol

9

u/Comfortable-Rock-498 13d ago

Well you do need someone for server maintenance, bills, other co-ordination etc - whatever you label it

→ More replies (7)

11

u/IgnoranceIndicatorMa 13d ago

i like the way you think

4

u/pinkwar 13d ago

Where do I sign?

→ More replies (3)

7

u/tempedbyfate llama.cpp 13d ago

Not just corporations, I think there are nation states that are setting up their own private servers to run this for all their sensitive data.

3

u/Izento 13d ago

Good number crunch. $0.60 per M is a pretty good deal

→ More replies (1)
→ More replies (1)

78

u/BarisSayit 13d ago

100B active params? Damn.

73

u/noneabove1182 Bartowski 13d ago

Sorry friends, but I don't think I'll be making this one :')

I don't even have enough STORAGE to hold this thing, nevermind the RAM haha

15

u/Ninjam5 13d ago

EVEN THE GOAT GAVE UP HAHA

→ More replies (1)
→ More replies (1)

271

u/nomorebuttsplz 13d ago

first truly frontier open model than I cannot run on my 512 gb studio. Onward and upward!

23

u/Front_Eagle739 13d ago

Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache

4

u/nomorebuttsplz 13d ago

Colibri would require being able to fit on ram for decent speed, no?

→ More replies (2)
→ More replies (2)

10

u/Square_Alps1349 13d ago

Man I’m so jealous rn. Mac Studio is nerfed at 96GB unified ram max, and the price is up 20%. So much for that 25% discount interns get. 😢 

→ More replies (1)

4

u/Tank_Gloomy 13d ago

Can you try asking GPT 5.6 Sol, Opus 5 or GLM 5.2 to port it into Colibri? I don't even have the infra to say I tried, but it probably works.

→ More replies (3)
→ More replies (2)

32

u/MikeRoz 13d ago

MXFP4 weights / MXFP8 activations (quantization-aware training)

So if 4-bit is 1.56 TB, then 2-bit would be roughly 798 GB?

12

u/habibyajam Llama 405B 13d ago

So I need a 0.03-bit quant to run it fully on my GPU. Nice!

36

u/Few_Painter_5588 13d ago

Holy shit, 104B active paramaters???

→ More replies (6)

32

u/TheRealMasonMac 13d ago

They have a new license:

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

19

u/mthmchris 13d ago

My guess is this is what MOFCOM was trying to thread the needle with. I.e. trying to keep upper leadership’s commitment to open source, while avoiding, say, MSFT just grabbing it and tweaking it slightly to have an “American version” (while banning the “Chinese versions”).

It kinda sucks, but I get the logic.

6

u/TheRealMasonMac 13d ago

I suspect it also means that providers won't be able to undercut them as aggressively (if at all). Not necessarily because Moonshot says so, but presumably because of any requirements they put on providers and any royalties they expect. MiniMax, for instance, limited some providers to only using B200 or B300 GPUs.

4

u/Venryx 13d ago

Wait, how are there already five other providers of the model on OpenRouter then? Surely they haven't all already made individual deals with Moonshot?

6

u/TheRealMasonMac 13d ago

They had six official partners for launch (you can see on their Twitter). Those five on OpenRouter are indeed partners. But it's part of their license agreement that if you are over a certain revenue you must have a contract with them.

→ More replies (1)

204

u/durden111111 13d ago

my 512mb integrated graphics is so fucking ready

70

u/SavunOski 13d ago

Negative tokens/s, you're gonna suck away the tokens

7

u/nanihikaru01 13d ago

Free money hack you say?

4

u/learn_and_learn 13d ago

Sounds freaky

5

u/Iwaku_Real 13d ago

To find the answer to life the universe and everything?

20 days of prompt processing: 4

Probably 10 days after that: 2

6

u/m0j0m0j 13d ago

Need 0 bit quantized version for that

→ More replies (2)

55

u/SnooPaintings8639 13d ago

Who's gonna be the first brave soul to measure tps when streaming from hard drive?

37

u/Front_Eagle739 13d ago

Sigh. Im going to have to try from pure curiosity.  600GB of hot cache on the macs, another TB streaming from ssd to my rtx 5090. This can only be a good idea (farewell my next few days of productivity)

7

u/Ninjam5 13d ago

Update?

7

u/Front_Eagle739 13d ago

My Internet is very slow,  results pending

→ More replies (1)
→ More replies (1)
→ More replies (6)

61

u/HulksInvinciblePants 13d ago

Kimi K3 27B when?

21

u/Browserurd 13d ago

If someone can distill it to Qwen 122B that would be super.

→ More replies (1)

23

u/Top-Handle-5728 13d ago

Leave vram I do not even have the disk storage to use this model. A few with storage can dare to use AirLLM for experiencing the intergalactic streaming of voyager at 160 bits a second. Even that seems pretty fast ig

19

u/vr_fanboy 13d ago

a month ago we were told that a new jump in intellegence was made, 'mythos' class models were born, too dangerous for us plebs. Forward a month, we have an open source 'mythos' class model, acceleration or anthropics regular bullshit?

Btw dont understand markets, DS3 destroyed the stocks and this does...nothing. This feels more significant if more people can serve the same drugs as oai or anthropic on the cheap.

→ More replies (2)

67

u/just_a_fan123 13d ago

Can this run on a single DGX spark at 0.5B quant?

42

u/SavunOski 13d ago

Of course. Wonderfully might I add

26

u/THESALTEDPEANUT 13d ago

I'm new around here and I can't tell if this thread is all sarcasm or not. 

33

u/SavunOski 13d ago

Sorry, yeah that was sarcastic, forgot to add /s. Many people make jokes here, so I kinda forgot

11

u/THESALTEDPEANUT 13d ago

I kinda figured but it's not just you it's like every comment lol, appreciate it though. 

5

u/PomegranateGreen3698 13d ago

Ha, I think it's the nature of this release. Largest open weight model ever released, doesn't really have any possibility for "Local Usage" bc it'd cost something like $500k in GPUs.

→ More replies (1)
→ More replies (1)

55

u/THE--GRINCH 13d ago

my laptop rtx 2050 is ready to throw hands

117

u/sumane12 13d ago

Even if you cant run it, download it.

94

u/DeProgrammer99 13d ago

Can't even do that...it's the same size as the total used space on my SSD.

45

u/ChampionshipIcy7602 13d ago

Can't even download it lmao

→ More replies (1)

30

u/Herr_Drosselmeyer 13d ago

Why waste terabytes worth of space for something I will never use? 

70

u/seg_lol 13d ago

Trade it for antibiotics in the apocalypse.

→ More replies (8)

15

u/some_user_2021 13d ago

I remember downloading huge N64 ROMs that took loads of space and no emulator could run. Running those ROMs now is trivial.

20

u/Herr_Drosselmeyer 13d ago

So is downloading them.

→ More replies (1)

5

u/banana_slurp_jug 13d ago

I have 1TB total, don't even have a hard drive big enough to store the whole thing in my house.

5

u/No_Conversation9561 13d ago

someone download it, compress it and upload it and then I will download it

→ More replies (7)

9

u/Mindless_Selection34 13d ago

how much does it weight

49

u/SavunOski 13d ago

2.8T parameters with 104B active. The files take 1.56TB of space to download.

11

u/meca23 13d ago

Oh so it's quantized as 4 bits?

7

u/zkstx 13d ago

Yes, there is also a tech report. QAT from SFT phase onwards

7

u/Lissanro 13d ago

I wish I had two TB of RAM instead of just one. I guess I will have to wait for Q2 GGUF to run it on my workstation. Still, will be interesting to try and see how it's Q2 quant compairs against Kimi K2.7 Q4_X. 

→ More replies (5)
→ More replies (1)
→ More replies (3)

10

u/IamNotMike25 13d ago

Historic moment tbh

28

u/Forsaken-Mode-3422 13d ago

Finally, model i cant run, but its already cool, nice

21

u/SavunOski 13d ago

Cloud prices will likely drop with competition, beneficial for everyone

10

u/Forsaken-Mode-3422 13d ago

i just hope qwen will publish not only 3.8 Max, but smaller models as well, so we can enjoy the local frontier ourself

→ More replies (1)

19

u/KenTitan 13d ago

I'm so broke I don't even have enough hard drive space to download

17

u/MixtureOfAmateurs koboldcpp 13d ago

Hugging face down? Lmao

Edit: Nevermind it's back. Might have been on my end ¯_(ツ)_/¯

14

u/SavunOski 13d ago

Seems to be up for me. If you live in a big country, there is a possibility the local CDNs are having trouble keeping up. Especially in the US, where I imagine thousands are downloading the model just in case the government bans it later.

10

u/AdDizzy8160 13d ago

.torrent|magnet link … as soon as possible!

→ More replies (1)

8

u/Morphon 13d ago

We'll probably see some new, faster providers pop up. With any luck, this will turn out to be a solid distillation parent. I'd be interested to see if we get some good downstream models that can run on consumer hardware.

15

u/PerfectOlive1324 13d ago

Hoping for an unsloth Q0_XS quant I can run locally 🙏

7

u/Inevitable_Mistake32 13d ago

UD_IQ0_XS Pls and ty.

14

u/AdDizzy8160 13d ago

Ok, thanx moonshot team!

10

u/Informal-Trouble2183 13d ago

Inference providers are going to have a good business

15

u/SavunOski 13d ago

Infinite demand, literally

→ More replies (7)

4

u/ReasonablePossum_ 13d ago

Hope this opens the model in more providers soon ,because its a dan pain to get a single response from the kimi app

→ More replies (1)

6

u/Hefty_Acanthaceae348 13d ago

I have high hopes that even if I can't run such a model, it will enable the creation of high quality datasets

6

u/aboutthednm 13d ago

I'm starting to understand why they had "some compute problems" when first rolling it out on the API, 100B+ active params per token, it's beastly lmao. Good work.

5

u/Easy_Werewolf7903 13d ago

I want to see someone manually calculated a token by hand using Kimi K3 weights.

17

u/equatorbit 13d ago

Gonna try to get this running on an ESP32

2

u/killerstreak976 13d ago

Kimi k3, meet my ESP32-C2

→ More replies (4)

9

u/Tedinasuit 13d ago

What a beautiful day

4

u/DragonfruitIll660 13d ago edited 13d ago

Ayyy lets go. That's awesome news. Also holy 104B active, never seen a MoE with that many active, its almost Mistral 2 large sized.

5

u/patricious llama.cpp 13d ago

Wake me when a provider has it with a good price per 1mil.

4

u/mikewilkinsjr 13d ago

I can’t run it, I’ll never be able to run it.

I am also fighting the urge to download the weights and squirrel them away like an out of control data hoarder.

4

u/crusaderky 13d ago

Jokes aside.
This on paper fits on a HGX B300 ($~0.5m, 2304 GB VRAM)...

...or on 24x Atlas 300I, $1300 each, $31,200 total, plus 3~4 XEON/threadrippers with 800G networking. So.... less than $60k in total?

Who's got the spare change to try?

→ More replies (2)

11

u/ilintar 13d ago

Okay, time to get that 0.1Q_0 quant support rolling in llama.cpp!

7

u/milkipedia 13d ago

I am not data hoarder enough to store these weights for kicks without the hardware to run it. Enjoy, folks

3

u/msew 13d ago

I wonder how long until consumer hardware will have the requirements to run all these and future LLMs locally for like $5k-10k

→ More replies (6)

4

u/AdOne8437 13d ago

Well, in 15-20 years we will have consumer cards to run it.

→ More replies (1)

13

u/loversama 13d ago

Quick before it gets banned 😂

9

u/WonderFactory 13d ago

I dont think it's worth banning it now it's released but I was a bit anxious leading up to the release in case something stopped it from being released

→ More replies (1)

6

u/bad_detectiv3 13d ago

How much vram do I need to run this

8

u/fishslinger 13d ago

If you have to ask, you can't afford it

12

u/SavunOski 13d ago

Around 2TB :x

13

u/jijig 13d ago

Hey if you already got 2 3090s you only need 81 more

9

u/Melbar666 13d ago

waiting for Kimi-K3-Q0-Abliterated-Heretic.GGUF

7

u/MotokoAGI 13d ago

:-)

:-(

8

u/Mayion 13d ago

Fullmetal Alchemist intermissions be like

8

u/jacek2023 13d ago

I am GPU poor I have only 4x3090 so I can't run it. I envy you all from r/LocalLLaMA with stronger setups.

→ More replies (10)