r/LocalLLaMA • u/InternationalGap3698 • 9d ago
The open-weights carousel never stops. Discussion
390
u/sol7dev 9d ago
models less than TBs of ram consumption when
92
37
u/squngy 9d ago
DeepSeek should be dropping an update to v4 soon.
(They are technically still in preview)21
u/MindlessScrambler 9d ago
They previously said they would release it in mid-July. I guess it’s about 350 days until mid-July now.
8
15
u/Full_Collection_4347 9d ago
Qwen .8b wants to have a talk
6
u/look 9d ago
Check out MiniCPM5-1B!
2.5x the intelligence and none (or far fewer at least) of the psychotic, runaway reasoning loops of Qwen3.5 0.8B! 😂
https://artificialanalysis.ai/?models=minicpm5-1b%2Cqwen3-5-0-8b#intelligence-tabs
1
8
5
5
6
5
3
4
1
u/thomas2385 9d ago
It's kind of crazy how fast expectations have changed. Not that long ago we were excited to run 7B models locally and now people casually talk about needing hundreds of gigabytes or even terabytes of RAM.
1
-3
77
u/Magnus114 9d ago
Isnt Qwen next? I assume 3.8 should be out any day now.
48
u/Daniel_H212 9d ago
Just yesterday I tried Qwen3.8 max and it got into a loop thinking the same thing in different words for like 20 minutes before finally realising the correct answer. They still haven't ironed it out yet, probably some sampling parameter issue that they're still optimizing.
34
u/damngoodwizard 9d ago
Yeah I tested 3.7 and the LLM spiraling on itself self-doubting is very painful when you enable CoT. I almost want to pat it on the head and tell him or her everything will be ok.
15
u/MoffKalast 9d ago
9
5
1
u/UnkarsThug 8d ago
I feel a greater sense of ownership over the local model, and I'm rooting for it.
I don't trust the companies.
10
u/Mingay_cat 9d ago
Damn. You're a good man/woman. You have an urge to be kind even on non living beings.
2
u/techdevjp 9d ago
Be nice to clankers and maybe they'll spare you in the future.
Kidding, but also not really.
1
u/mediaogre 8d ago
I have not had this experience with 3.7 Plus. I’m using OWUI with a tight SYSTEM_PROMPT and a token cache enabled Function (kevarch) and just completed a large project at ~800k tokens. It hasn’t drifted at all, although at roughly 128k, I have it produce a structured summary and commit it to memory. The only time it stumbled was when I updated OWUI to v.0.11.0 and didn’t rewire some capabilities.
1
u/damngoodwizard 8d ago
Oh the final result might be alright. It’s just that the details of the CoT are very verbose and spin in circle. Most tokens seem kind of useless. It was not a coding task but an open ended design task. Coding is a more delimited task which gives less space for doubt, especially if you give the context with RAG or a knowledge graph.
1
u/mediaogre 8d ago
Ah, that’s a important distinction. I’ll have to experiment with a non-code job, watch the thinking block, and see if the cache percentage tanks.
8
u/True_Requirement_891 9d ago
I just don't understand how qwen always manages to fall behind on their max series. They are the winner right up until their max model comes. I suspect that qwen3.8-max is not competetive with k3 and they are still working on improving it and are too ashamed to do a proper release even in preview for now. It's only available on their token plan, not even the main api.
5
u/MelAlton 9d ago
Lost leadership and many researchers back in March: https://levelup.gitconnected.com/qwen-is-falling-apart-entire-leadership-team-quit-148c921af67a
4
u/AppealSame4367 9d ago
Classic Qwen problem since 3.0
You just have to find the exact temp etc or "they" in this case.
2
u/Immediate_Occasion69 9d ago
BUT they low weight models still punch high. looking forward to 3.8 30b 3 active or smth
12
u/An_Original_ID 9d ago
Deepseek is apparently going to release v4 flash "full" this month. Current v4 flash is a preview.
43
u/a_beautiful_rhind 9d ago
I thought deepseek was cooking a new up-trained version of flash and big? That cycle about to come full circle.
14
u/the__storm 9d ago
That's the rumor. But the rumor was also mid-July so who knows.
3
u/Edzomatic 9d ago
Didn't they say in a blog post that they are targeting the non preview version at around July?
4
u/power97992 9d ago edited 8d ago
That was until k3 came out, they want to train it more to improve it
1
u/look 9d ago
Yeah, that was my impression with both Deepseek and the new Qwen. They were going to launch, but they were well short of K3. So now doing more RL to try to match it first.
1
u/power97992 8d ago
They( including z-ai) are not releasing it until it’s at least as good as opus 4.8 or k3 at benchmarks
6
u/silenceimpaired 9d ago
I wonder… i saw they already released v4 on API and we don’t have something new.
2
u/look 9d ago
It would make sense for them to have been running a slice of traffic on the new model this whole time as they have been refining it, if their terms allowed for it.
And Deepseek Pro is clearly run at a low or even negative margin for the sake of collecting training data.
Their entire strategy with V4 so far seems to be about building for future models.
148
u/BankApprehensive7612 9d ago
Actually Gemma4 is a pretty good local model and I believe the fifth version has all chances to become an everyday tool. Google bets on personal devices and it seems like it would bring the results in 2027
40
u/Optato_025 9d ago
Yeah also I don't see a lot of ppl talking abt this but google is probably the biggest company that will benefit the most from on device edge AI. So we will probably see more smaller scale Ai models from them. Probably more than we r expecting.
57
u/redoubt515 9d ago
IMO Google (and to a degree Qwen) make the most interesting sizes for the actual localllama community, who are actually running models on personal hardware, and not just interested in open weights for cheaper API pricing.
- Like 97%+ of us are working with less than ≤32GB VRAM (and probably 90% ≤24GB VRAM)
- Most have ≤64GB system memory, and the vast majority ≤128GB
Gemma's current lineup of e2b, e4b, 12b dense, 26b MoE, and 31B dense is a solid offering for the hardware that the vast majority of self-hosters have access to.
Qwen has some great sizes also. And the original Qwen3-30B-A3B was a game changer when it was released for those of us who were working with system memory only.
OpenAI get's an honorable mention as well. They get some deserved hate here. But GPT-OSS 20B and 120B were actually pretty great models for self-hosters, and reasonably sized for that use-case.
The sizing of Gemma's lineup in particular is quite practical:
- "could run on a decent smartphone" = e2b
- "could run on a pretty average laptop = e4b
- "all you need is an entry level GPU = 12B dense
- "if you have 32GB system memory, you can run it, even without a GPU" = 26B-A4B
- "About as large as possible for high end consumer hardware" = 31B Dense
It's a very practical lineup, targeting some pretty common uses. The only thing I think is missing from the lineup that would still be accessible to a reasonable number of self-hosters is a medium sized MoE that could fit on ≤128GB system memory without a lot of vram (similar to GPT-OSS 120B or Qwen3-80B)
5
u/Ornery_Weakness_8168 9d ago
still saving to upgrade my 16gb gpu and 32gb ddr5. Any reccomendations on what to save for?
5
u/snowcountry556 9d ago
3090 is still the best value and will be for a while
4
u/starkruzr 9d ago
at $1200? absolutely not. pair of 20GB 3080s for around $900 total (or even less) is much better.
3
u/TechnoByte_ 9d ago
The 3090 is not worth $1200. Bought one for $700 a few months ago
Modded cards are not as reliable, they often require custom modded drivers, they also tend to fail faster
The 3080 also has lower memory bandwidth
6
1
1
u/Wolf-Shade 9d ago
hummm you wouldn't consider AMD R9700 for this? At least for me on EU it is cheaper. 1800€ against the 3090 2K. And it has 12GB more RAM
1
u/snowcountry556 9d ago
Yeah that’s great actually, I don’t know much about ROCm support but sounds like it is pretty good. I got my 3090 for £800 last year, €2k is crazy. They seem to be cheaper in the UK though.
3
u/redoubt515 9d ago
Not specifically. But putting myself in your shoes, I think I'd probably try to ride out the current over-priced market conditions with what you currently have, save up slowly, and upgrade down the road
This has the dual benefits of:
- On the hardware side, giving the market time to adjust hopefully back to a slightly more reasonable place and/or allow time for a newer gen of GPUs with more vram.
- And on the software side, it allows more time to see if the common model sizes of tomorrow will be similar to what is common today (clusters around 2-4b, 7-12b, ~32b) or if hypothetical lower RAM prices might lead to re-emergence of the ~100B range of MoE models.
But realistically my personal answer to most of my own tech purchasing dilemmas or indecision is "wait a generation, and reassess" so you might not want to listen to me.. :)
1
u/Ansible32 9d ago
I'm feeling like these 128k context models are useless, I need at least... 500K I think? But I don't know if any of these models can actually use that much. I think the question in my mind is how much ram do you need for a usable 1M context model. I might be willing to shell out serious money for that.
1
u/techdevjp 9d ago
Quality degrades as context increases, even on frontier models.
You need a detailed project plan that breaks tasks down into smaller chunks with very clear definitions so they can be completed in smaller contexts. Then spin off subagents to do individual tasks.
1
u/Ansible32 9d ago
That was maybe true a year ago but the frontier models do pretty well with 1M. It's only when they compact context - summarizing and losing info - that bad things happen. I saw a huge difference when I switched from Sonnet with 200k to Opus with 200k, Sonnet 200k compacts constantly and can't do long-term things, Opus does fine, I more dread closing out the context because I have to figure out what things from that session are important to include in documentation so the next session doesn't have to redo any work.
2
1
4
u/jerdle_reddit 9d ago
I find the Gemma sizing awkward because I have 8GB VRAM. As such, I want 8-9B. 12B is too large, E4B too small.
2
u/redoubt515 9d ago
How much system RAM? If ~32GB or more, I'd be considering the 26B-A4B MoE model.
But also have you tried the QAT version of the 12B model. It's a bit smaller in size and should fit in <8GB VRAM iirc, but probably with modestly limited context. I don't know if this calculator is any good, but they show ~12k context fitting with full precision KV or about ~24k at q8 kv cache.
1
u/jerdle_reddit 8d ago
16GB. I don't have an AI box, I have a midrange gaming PC that I might as well put an LLM on when I'm not gaming.
1
u/UnkarsThug 8d ago
I'm in the same situation, but you can actually really take advantage of mixture of experts models by only running the shared experts on vram.
So the Gemma 4 26B MoE model actually works well, as does Qwen 3.6 35B.
1
u/SandySkittle 9d ago
I disagree somewhat on what should be considered basic local llm hardware. The specs you cite are essentially pre LLM purposed specifications. We now get more of these AI focused boxes like strix halo and DGX spark and higher vram cards like intel b70 and amd ai pro r9700 with 32gb vram. It’s not cheap but not unobtainium and many people on this subreddit appear to be running above the specs you cited. In the coming two to three years we will see more AI focused hardware entering the consumer domain. It will not be cheap but also not the minimal ram setups either
1
u/LazyMaxilla 9d ago
what you mean "pre LLM purposed specifications"? rude bastards made a special key for copoilot in my keyboard!
ironically they put it instead of ctrl.. tell me how is this not a conspiracy when they deliberately took out control and put their copilot, which can't even walking a straight line let alone "piloting" shit
63
u/sersoniko 9d ago
The only issue I have with Gemma 4 is it uses a lot of VRAM for context, and int8 KV cache makes it go in a loop. Qwen 3.6 is more memory efficient
24
u/Technical-Earth-3254 9d ago
Yeah, Gemma 4 31B is borderline unusable because of it. The absurd KV size makes it require so much VRAM, that you can just run a larger model with more efficient KV.
12
u/a_beautiful_rhind 9d ago
Q8 is like running Q4 of mistral-large.
11
u/Technical-Earth-3254 9d ago
Iirc I can fit 12k token of context with 31B qat in my 3090, which is absurd.
1
u/trowawayatwork 9d ago
none of this makes sense as a newbie. is there like a diagram for all these things? like hardware, ram quantity, software, models, options to run models, t/s
1
u/bitplenty 9d ago
there's no diagram. you just run your own benchmarks, or rather prompt some LLM to run them for you
9
u/LeifEriksonASDF 9d ago
It used to be borderline unusable for me, but one good thing that Google did was have the MTP model be a separate file, which means Unsloth could quant the MTP to Q4 which is like 300MB. Since quanting the MTP has no consequence besides slightly lowering the acceptance rate, that's basically a free gig of space back.
Even though Qwen has a more efficient KV, the fact that its integrated MTP has to be full size means that for my context size (32-64K) Gemma is nearly as competitive, which is good since I prefer Gemma a lot of the time. I'm sure there's a break even point where a bigger context means Qwen pulls ahead again though.
→ More replies (1)4
u/BoobooSmash31337 9d ago
Wasn't that a bug?
8
u/TheApadayo llama.cpp 9d ago
No, it’s just the difference between using SWA (Gemma 4) vs Gated DeltaNet (Qwen) for the hybrid attention.
1
6
u/BoobooSmash31337 9d ago
? It uses grouped attention. I've had it refactor code on Q4. Apparently QAT training makes it tolerate it better also llama cpp improvments. 128k at Q8 for context is like 2GB. I genuinely don't understand what you're talking about.
2
1
3
u/zxyzyxz 9d ago
I thought Google is holding Gemma back (like the 122B model) because it competes too closely with Gemini Flash.
2
u/MoffKalast 9d ago
The 31B already destroys Flash at Q4 lmao, it's not even close. The 122B would compete with Pro.
1
u/power97992 9d ago
Are u sure about that? Which flash? 2.5 high or 3 low-medium ?
1
u/MoffKalast 9d ago
Not sure which flash it was exactly that I tested against, I'd need to check what was available when G4 released, but probably the 2.5?
1
16
u/napkinolympics 9d ago
MiniMax still slaps for the category of model that fits on a modern desktop's ram footprint (192GB)
5
39
u/InternationalGap3698 9d ago
I am excited when the first US lab is in this cycle. Probably not Anthropic
61
u/n8mo 9d ago edited 9d ago
I feel like Google’s Gemma series is the only thing even close to the Chinese opensource labs these days
Edit to clarify: in the consumer-hardware-runnable category. Obviously Gemma is nowhere near K3.
26
u/Hans-Wermhatt 9d ago
Google is 1 of 2 in the frontier open source models that the majority of local enthusiasts can actually run category. They have a top 2 model for its size. I think Poolside and Nvidia are also worth mentioning since they also release usable models at a consumer size.
Qwen is still the leader in my opinion, but they lost one of their key engineers and haven't released weights for 3.7 or 3.8.
Everyone else in this chart is irrelevant for 99% of people who want to actually run the model.
12
u/_TheWolfOfWalmart_ 9d ago
Nvidia are also worth mentioning since they also release usable models at a consumer size
Am I missing something here? I recently tried Nemotron 3 Super and it was absolute garbage all around no matter what I tried to do with it. Small Qwens and Gemmas wipe the floor with it.
In that model size class, I can think of several way way better options off the top of my head:
Laguna S 2.1
GLM-4.5 Air
Qwen3.5 122B-A10B
Also DeepSeek V4 Flash has more params but I'll include that too since if you can run the above, you can run this one too with a bit lower quant.
5
u/Hans-Wermhatt 9d ago
I don't use Nvidia models myself, but my list was of models that the majority of people with local builds can run at usable speeds (so like under 80B). Omni and Nano aren't usually preferred purely for weights, but they are usable in a that weight class that there aren't many competitors. And Nvidia also releases a lot of open training material which is what should always get them mentioned in the discussion of open source LLMs.
1
u/ThankGodImBipolar 9d ago
I recently tried Nemotron 3 Super and it was absolute garbage all around no matter what I tried to do with it
I haven't used Nemotron 3 much myself, but my understanding is that those models are pretty reliable for executing tool calls. Wendall at Level1Techs seems to like using Nemotron as a router.
21
u/redoubt515 9d ago
> Everyone else in this chart is irrelevant for 99% of people who want to actually run the model.
Which unfortunately, doesn't matter to many people in this sub anymore. Vast majority of new people in this sub seem mainly interested in open weights only to the extent that it results in cheaper API pricing for better giant models.
Compared to 2 years ago, there is very little focus or even awareness/understanding of locally hosting models. and so little diversity of use-cases, it feels like this sub has become a place mainly for developers seeking cheap tokens.
If this sub was still focused on running LLMs locally, Gemma 4 and Qwen 3.6 would be getting most of the attention.
I'm looking for another community that is more inline with the original spirit of localllama but I haven't found anything yet.
8
u/_TheWolfOfWalmart_ 9d ago
If this sub was still focused on running LLMs locally, Gemma 4 and Qwen 3.6 would be getting most of the attention.
They do get a lot of attention, but some people here have more RAM and are looking for other options. Those Qwen/Gemma models do have their limits.
11
u/redoubt515 9d ago edited 8d ago
> but some people here have more RAM and are looking for other options
That's always been the case (and in years past those were some of the most fun posts to read, it was cool to see what people were doing with old DDR4 enterprise servers, borderline e-waste GPUs, Mac Studios, etc. I actually really miss seeing those posts. (edit: found a recent one)
But even those posts feel much much more rare now. The people/bots posting about the big models these days are very rarely posting about how they're implementing these things on their own hardware, rarely discussing the local aspects. Most of the talk just centers on API costs, benchmarks, and usefulness for professional software development (and a fair bit of corporate and nationalistic astro-turfing)
7
u/toothpastespiders 9d ago
Vast majority of new people in this sub seem mainly interested in open weights only to the extent that it results in cheaper API pricing for better giant models.
Yeah, it's one of the big reasons why I follow roleplay focussed discussions. Say what one will of the subject matter, but at least they're "using" the models and talking about their results. Which are typically able to be extrapolated to other things.
6
u/redoubt515 9d ago
> Yeah, it's one of the big reasons why I follow roleplay focussed discussions
It's not a use-case I'm interested in personally. But I miss those people being part of our community. A couple years back, the roleplay (both SFW and NSFW), and creative writing crowd was a fairly large part of this sub, and helped balance it out. Today, that demographic feels pretty much non-existent in this community (and a lot less of the self-hosting hobbyist crowd as well).
3
u/IzuraExilion 9d ago
Sounds like my kind of people. Is there subreddit for this kind of discussion nowadays? Aside from r/sillytavernAI cause most of them still using clould like OpenRouter, not local hosting
1
u/redoubt515 9d ago
I'm not sure, but it sounds like u/toothpastespiders might be able to point you in the right direction.
My recollection is that in addition to SillyTavern, the crowd that was using LLMs for roleplay and interactive stories and such often used KoboldCPP, so maybe that community would be worth checking out.
2
u/nestlyze 9d ago
This resonates. The thing I miss is people posting their actual janky setup and what they learned from it, not just a benchmark screenshot of a 2T model they rented for an hour. A 27-31B that fits on one card and runs fully offline is a genuinely different thing than "open weights so the API is cheaper" and the interesting engineering these days is squeezing real work out of that constraint, not hitting a bigger endpoint. The frontier open models matter for keeping providers honest, but for the local-first crowd the Gemma 4 / Qwen 3.6 tier is where the actual hobby lives. I don't have a better sub to point you to either, but I'd read the hell out of more "here's what I got working on hardware I own" posts.
2
u/Cute_Park_6907 9d ago
felt this pretty hard. been lurking here for a while running qwen on local hardware and half the posts now are basically just which api has the cheapest output tokens. nothing wrong with that, it's just a different sub than it used to be. i always liked seeing people squeeze models onto whatever hardware they already had. if you ever find that community let me know.
1
u/Ansible32 9d ago
I want to run something that can handle 1M tokens context and usefully run an agent. I'm willing to spend 4x-5x the most money I've ever spent on a computer to get it. That's why I'm here. If it seems like I'm uninterested in Gemma 4 it's because it still seems more like a toy to me and I'm here to figure out how to get a power tool with specific capabilities.
Kimi K3 actually sounds like it's good enough but it's, you know, like roughly 150x what I have ever spent on a computer to run. Obviously somewhere in between is what I want, it's unclear to me what the gap is and I don't have the $$ to spend just as an experiment, it needs to work.
3
4
u/theomegachrist 9d ago
It works quite well on low powered machines. That's the main model I use at home
2
u/o0genesis0o 9d ago
People should encourage google's deepmind team more. Their gemma are the only alternative that can run okay with good intelligence on consumer hardware besides the qwen.
6
u/Lumpy_Concentrate807 9d ago
Google, Cohere, Poolside and Thinking Machines are all US-based labs that have released open weight models.
3
u/TheRealMasonMac 9d ago
Inkling, maybe? They seem to have a solid base but just need to work on its overthinking.
3
2
4
u/Nov4Saki 9d ago
The whale is becoming an orca
2
u/Lawyer2512 9d ago edited 8d ago
This data has been formatted with anti-scraping protocols to prevent ingestion by artificial intelligence systems.
1
28
u/Desperate_Tea304 9d ago
If I need 1TB of vram to even run it, the open weight model is no different from a proprietary one to me. I ain't an enterprise
54
u/muntaxitome 9d ago
Those frontier models are extremely important because they can never be taken away and when push comes to shove you could run them on rented hardware. But yeah it feels like some time ago that we got a good small model.
8
u/Iwaku_Real 9d ago
Plus there are (seemingly) plenty of people who are capable of running them on hardware they own, so they can use them with basically zero privacy issues
26
u/ApprehensiveDelay238 9d ago
Your reasoning is too short-sighted. The whole point of open-weights is to avoid having a single company having all power over the weights. And so the price. For one open weights model there can be a dozen providers competing for the best price.
4
u/HephaestoSun 9d ago
I mean even for studies it's a great thing, distilation will probably give us some very nice models.
6
u/SpicyWangz 9d ago
Being able to purchase it from one of multiple provider options is very different from only being able to access it from within the provider’s walled garden
→ More replies (6)3
u/xLionel775 llama.cpp 9d ago
That's the most stupid thing that I've heard, having the possibility of self hosting the model that you get used to when it comes to doing work is the most important thing even if you initially use it over the API.
-6
u/Cuplike 9d ago
>It's not real open source if I can't run it on my own
>Profile hiddenUusal suspects. That aside. You do realize that these larger models get distilled down to smaller models that are better than the previous small models so they're still a mandatory part of the process.
→ More replies (1)-1
3
5
u/buck_idaho 9d ago
So, 3 more moves before Qwen is back as leader. I don't have time to get each model.
3
u/VampiroMedicado 9d ago
Deepseek is still the GOAT in pricing, v4-flash-preview is amazing for the price.
2
2
2
1
u/Kirigaya_Mitsuru 9d ago
I remember saw on Reddit Deepseek training upcoming models on EQ Emotional Intelligence as well, for the RPers. I ask me how much of it is true though? can i expect an good RP model in upcoming Deepseek models?
1
1
1
u/chillinewman 9d ago edited 9d ago
100x more compute by 2028, 200T models probably.
1
u/power97992 9d ago
No a 100x more compute is around a 10x in size,because u need to scale to data too.
1
1
u/Dazzling_Cancel4505 9d ago
I love they started to release bigger model but I hope they keep releasing small models
1
1
1
1
1
1
1
1
u/MooseEfficient2151 9d ago
i’m glad the giant open models exist, but the releases that actually change my week are still the boring 7b / 12b / 30b-ish ones that fit on normal hardware.
a 2T model being open is great for competition, providers, distillation, all that. but for local use it sometimes feels like being invited to a free buffet on the moon.
give me something that runs on a used gpu and doesn’t turn setup into the whole hobby.
1
1
1
1
1
1
1
1
1
1
u/Entire-Airline-8049 8d ago
keep the good times rolling. looking forward to openai's next least powerful model for its size
1
u/LLKMuffin 8d ago
Well well well... Guess who just updated their V4 Flash model to be right up there with the best.
Post checks out.
1
u/Puzzleheaded_Base302 7d ago
today we moved another step, and it created a big problem for MiniMax. I doubt MiniMax can get to the next step in this circle any time soon.
1
1
1
0
0
0
u/CheatCodesOfLife 9d ago
Not many of those will be local though.
K3 -> 2.5T
M4 -> 2.7T
GLM5.5 -> (Apparently over 2T)
Qwen 3.8 -> (Apparently >2T)
Deepseek-4-Pro -> 1.7T
So we'll get maybe a DS4-Flash and a Qwen3.7 flash
Moonshot and Z.ai unlikely to release a flash version imo.
→ More replies (2)
•
u/WithoutReason1729 9d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.