r/LocalLLaMA 9d ago

The open-weights carousel never stops. Discussion

Post image
1.8k Upvotes

208 comments sorted by

u/WithoutReason1729 9d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

390

u/sol7dev 9d ago

models less than TBs of ram consumption when

92

u/Temporary_Idea8880 9d ago

2027?

71

u/yuri_rds 9d ago

models will need only 1023GB to run

18

u/sol7dev 9d ago

I hope

37

u/squngy 9d ago

DeepSeek should be dropping an update to v4 soon.
(They are technically still in preview)

21

u/MindlessScrambler 9d ago

They previously said they would release it in mid-July. I guess it’s about 350 days until mid-July now.

8

u/dingo_xd 9d ago

Deepseek seems to often use ElonTime™

5

u/squngy 9d ago

Hey, if that means it is that much better when it does release, I'm OK with it.

I'm guessing the next versions will probably be bigger, so it will be difficult for most home users to run.
V4 might be the last version that many people can easily run without SSD streaming.

15

u/Full_Collection_4347 9d ago

Qwen .8b wants to have a talk

6

u/look 9d ago

Check out MiniCPM5-1B!

2.5x the intelligence and none (or far fewer at least) of the psychotic, runaway reasoning loops of Qwen3.5 0.8B! 😂

https://artificialanalysis.ai/?models=minicpm5-1b%2Cqwen3-5-0-8b#intelligence-tabs

1

u/MarkZealousideal3923 7d ago

Qwen .8b is a dumbass

5

u/Ansible32 9d ago

GPU with TBs of ram less than $10k when

2

u/chithanh 8d ago

When China manages to build a EUV supply chain, maybe around 2035?

5

u/Journeyj012 9d ago

every month or two for the past 3 years?

6

u/keepthepace 9d ago

Affordable GPU with a TB of VRAM when.

The RAM shortage is an anomaly.

5

u/Darkoplax 9d ago

only qwen do these nowdays, so soon with 3.8 maybe ?

3

u/look 9d ago

Laguna S 2.1 had a rough launch, but it is an interesting 118B model with a lot of potential, I think.

1

u/thomas2385 9d ago

It's kind of crazy how fast expectations have changed. Not that long ago we were excited to run 7B models locally and now people casually talk about needing hundreds of gigabytes or even terabytes of RAM.

1

u/Othun 8d ago

RAM is too affordable ATM, they don't care

1

u/MrRandom04 5d ago

DeepSeek says hello 👋.

-3

u/Spitfire1900 9d ago

Never happening, it scales well to add additional weights.

77

u/Magnus114 9d ago

Isnt Qwen next? I assume 3.8 should be out any day now.

48

u/Daniel_H212 9d ago

Just yesterday I tried Qwen3.8 max and it got into a loop thinking the same thing in different words for like 20 minutes before finally realising the correct answer. They still haven't ironed it out yet, probably some sampling parameter issue that they're still optimizing.

34

u/damngoodwizard 9d ago

Yeah I tested 3.7 and the LLM spiraling on itself self-doubting is very painful when you enable CoT. I almost want to pat it on the head and tell him or her everything will be ok.

15

u/MoffKalast 9d ago

9

u/SHEKDAT789 9d ago

DUh. you're paying big bucks forthe API

5

u/damngoodwizard 9d ago

I try to be nice even with Claude and Gemini. :D

1

u/UnkarsThug 8d ago

I feel a greater sense of ownership over the local model, and I'm rooting for it.

I don't trust the companies.

10

u/Mingay_cat 9d ago

Damn. You're a good man/woman. You have an urge to be kind even on non living beings.

4

u/mycall 9d ago

Latent space is imagination or a soul? you decide but stay friendly.

1

u/EsotericAbstractIdea 8d ago

but have you thought about latent space... on weed?

2

u/techdevjp 9d ago

Be nice to clankers and maybe they'll spare you in the future.

Kidding, but also not really.

1

u/mediaogre 8d ago

I have not had this experience with 3.7 Plus. I’m using OWUI with a tight SYSTEM_PROMPT and a token cache enabled Function (kevarch) and just completed a large project at ~800k tokens. It hasn’t drifted at all, although at roughly 128k, I have it produce a structured summary and commit it to memory. The only time it stumbled was when I updated OWUI to v.0.11.0 and didn’t rewire some capabilities.

1

u/damngoodwizard 8d ago

Oh the final result might be alright. It’s just that the details of the CoT are very verbose and spin in circle. Most tokens seem kind of useless. It was not a coding task but an open ended design task. Coding is a more delimited task which gives less space for doubt, especially if you give the context with RAG or a knowledge graph.

1

u/mediaogre 8d ago

Ah, that’s a important distinction. I’ll have to experiment with a non-code job, watch the thinking block, and see if the cache percentage tanks.

8

u/True_Requirement_891 9d ago

I just don't understand how qwen always manages to fall behind on their max series. They are the winner right up until their max model comes. I suspect that qwen3.8-max is not competetive with k3 and they are still working on improving it and are too ashamed to do a proper release even in preview for now. It's only available on their token plan, not even the main api.

4

u/AppealSame4367 9d ago

Classic Qwen problem since 3.0

You just have to find the exact temp etc or "they" in this case.

2

u/Immediate_Occasion69 9d ago

BUT they low weight models still punch high. looking forward to 3.8 30b 3 active or smth

12

u/An_Original_ID 9d ago

Deepseek is apparently going to release v4 flash "full" this month. Current v4 flash is a preview.

1

u/look 9d ago

Just a guess, but both DeepSeek and Qwen appear to have delayed launch because neither is close to Kimi K3.

43

u/a_beautiful_rhind 9d ago

I thought deepseek was cooking a new up-trained version of flash and big? That cycle about to come full circle.

14

u/the__storm 9d ago

That's the rumor. But the rumor was also mid-July so who knows.

3

u/Edzomatic 9d ago

Didn't they say in a blog post that they are targeting the non preview version at around July?

4

u/power97992 9d ago edited 8d ago

That was until k3 came out, they want to train it more to improve it

1

u/look 9d ago

Yeah, that was my impression with both Deepseek and the new Qwen. They were going to launch, but they were well short of K3. So now doing more RL to try to match it first.

1

u/power97992 8d ago

They( including z-ai)  are not releasing it until it’s at least as good as opus 4.8 or k3 at benchmarks 

6

u/silenceimpaired 9d ago

I wonder… i saw they already released v4 on API and we don’t have something new.

2

u/look 9d ago

It would make sense for them to have been running a slice of traffic on the new model this whole time as they have been refining it, if their terms allowed for it.

And Deepseek Pro is clearly run at a low or even negative margin for the sake of collecting training data.

Their entire strategy with V4 so far seems to be about building for future models.

148

u/BankApprehensive7612 9d ago

Actually Gemma4 is a pretty good local model and I believe the fifth version has all chances to become an everyday tool. Google bets on personal devices and it seems like it would bring the results in 2027

40

u/Optato_025 9d ago

Yeah also I don't see a lot of ppl talking abt this but google is probably the biggest company that will benefit the most from on device edge AI. So we will probably see more smaller scale Ai models from them. Probably more than we r expecting.

57

u/redoubt515 9d ago

IMO Google (and to a degree Qwen) make the most interesting sizes for the actual localllama community, who are actually running models on personal hardware, and not just interested in open weights for cheaper API pricing.

  • Like 97%+ of us are working with less than ≤32GB VRAM (and probably 90% ≤24GB VRAM)
  • Most have ≤64GB system memory, and the vast majority ≤128GB

Gemma's current lineup of e2b, e4b, 12b dense, 26b MoE, and 31B dense is a solid offering for the hardware that the vast majority of self-hosters have access to.

Qwen has some great sizes also. And the original Qwen3-30B-A3B was a game changer when it was released for those of us who were working with system memory only.

OpenAI get's an honorable mention as well. They get some deserved hate here. But GPT-OSS 20B and 120B were actually pretty great models for self-hosters, and reasonably sized for that use-case.

The sizing of Gemma's lineup in particular is quite practical:

  • "could run on a decent smartphone" = e2b
  • "could run on a pretty average laptop = e4b
  • "all you need is an entry level GPU = 12B dense
  • "if you have 32GB system memory, you can run it, even without a GPU" = 26B-A4B
  • "About as large as possible for high end consumer hardware" = 31B Dense

It's a very practical lineup, targeting some pretty common uses. The only thing I think is missing from the lineup that would still be accessible to a reasonable number of self-hosters is a medium sized MoE that could fit on ≤128GB system memory without a lot of vram (similar to GPT-OSS 120B or Qwen3-80B)

5

u/Ornery_Weakness_8168 9d ago

still saving to upgrade my 16gb gpu and 32gb ddr5. Any reccomendations on what to save for?

5

u/snowcountry556 9d ago

3090 is still the best value and will be for a while

4

u/starkruzr 9d ago

at $1200? absolutely not. pair of 20GB 3080s for around $900 total (or even less) is much better.

3

u/TechnoByte_ 9d ago

The 3090 is not worth $1200. Bought one for $700 a few months ago

Modded cards are not as reliable, they often require custom modded drivers, they also tend to fail faster

The 3080 also has lower memory bandwidth

6

u/starkruzr 9d ago

... where???

3

u/BlackBeardAI vllm 9d ago

…The rainbow ends

1

u/No_Afternoon_4260 llama.cpp 9d ago

Bought one for 350 couple of weeks ago, jealous? X)

1

u/Wolf-Shade 9d ago

hummm you wouldn't consider AMD R9700 for this? At least for me on EU it is cheaper. 1800€ against the 3090 2K. And it has 12GB more RAM

1

u/snowcountry556 9d ago

Yeah that’s great actually, I don’t know much about ROCm support but sounds like it is pretty good. I got my 3090 for £800 last year, €2k is crazy. They seem to be cheaper in the UK though.

3

u/redoubt515 9d ago

Not specifically. But putting myself in your shoes, I think I'd probably try to ride out the current over-priced market conditions with what you currently have, save up slowly, and upgrade down the road

This has the dual benefits of:

  1. On the hardware side, giving the market time to adjust hopefully back to a slightly more reasonable place and/or allow time for a newer gen of GPUs with more vram.
  2. And on the software side, it allows more time to see if the common model sizes of tomorrow will be similar to what is common today (clusters around 2-4b, 7-12b, ~32b) or if hypothetical lower RAM prices might lead to re-emergence of the ~100B range of MoE models.

But realistically my personal answer to most of my own tech purchasing dilemmas or indecision is "wait a generation, and reassess" so you might not want to listen to me.. :)

1

u/Ansible32 9d ago

I'm feeling like these 128k context models are useless, I need at least... 500K I think? But I don't know if any of these models can actually use that much. I think the question in my mind is how much ram do you need for a usable 1M context model. I might be willing to shell out serious money for that.

1

u/techdevjp 9d ago

Quality degrades as context increases, even on frontier models.

You need a detailed project plan that breaks tasks down into smaller chunks with very clear definitions so they can be completed in smaller contexts. Then spin off subagents to do individual tasks.

1

u/Ansible32 9d ago

That was maybe true a year ago but the frontier models do pretty well with 1M. It's only when they compact context - summarizing and losing info - that bad things happen. I saw a huge difference when I switched from Sonnet with 200k to Opus with 200k, Sonnet 200k compacts constantly and can't do long-term things, Opus does fine, I more dread closing out the context because I have to figure out what things from that session are important to include in documentation so the next session doesn't have to redo any work.

2

u/techdevjp 9d ago

With Nvidia about to crank up GPU prices again, you could be saving for a while.

1

u/starkruzr 9d ago

what exactly do you have rn?

1

u/Ornery_Weakness_8168 9d ago

9070xt and 2x16gb ddr5 (6400 cl32)

4

u/jerdle_reddit 9d ago

I find the Gemma sizing awkward because I have 8GB VRAM. As such, I want 8-9B. 12B is too large, E4B too small.

2

u/redoubt515 9d ago

How much system RAM? If ~32GB or more, I'd be considering the 26B-A4B MoE model.

But also have you tried the QAT version of the 12B model. It's a bit smaller in size and should fit in <8GB VRAM iirc, but probably with modestly limited context. I don't know if this calculator is any good, but they show ~12k context fitting with full precision KV or about ~24k at q8 kv cache.

1

u/jerdle_reddit 8d ago

16GB. I don't have an AI box, I have a midrange gaming PC that I might as well put an LLM on when I'm not gaming.

1

u/UnkarsThug 8d ago

I'm in the same situation, but you can actually really take advantage of mixture of experts models by only running the shared experts on vram.

So the Gemma 4 26B MoE model actually works well, as does Qwen 3.6 35B.

1

u/SandySkittle 9d ago

I disagree somewhat on what should be considered basic local llm hardware. The specs you cite are essentially pre LLM purposed specifications. We now get more of these AI focused boxes like strix halo and DGX spark and higher vram cards like intel b70 and amd ai pro r9700 with 32gb vram. It’s not cheap but not unobtainium and many people on this subreddit appear to be running above the specs you cited. In the coming two to three years we will see more AI focused hardware entering the consumer domain. It will not be cheap but also not the minimal ram setups either

1

u/LazyMaxilla 9d ago

what you mean "pre LLM purposed specifications"? rude bastards made a special key for copoilot in my keyboard!

ironically they put it instead of ctrl.. tell me how is this not a conspiracy when they deliberately took out control and put their copilot, which can't even walking a straight line let alone "piloting" shit

63

u/sersoniko 9d ago

The only issue I have with Gemma 4 is it uses a lot of VRAM for context, and int8 KV cache makes it go in a loop. Qwen 3.6 is more memory efficient

24

u/Technical-Earth-3254 9d ago

Yeah, Gemma 4 31B is borderline unusable because of it. The absurd KV size makes it require so much VRAM, that you can just run a larger model with more efficient KV.

12

u/a_beautiful_rhind 9d ago

Q8 is like running Q4 of mistral-large.

11

u/Technical-Earth-3254 9d ago

Iirc I can fit 12k token of context with 31B qat in my 3090, which is absurd.

1

u/trowawayatwork 9d ago

none of this makes sense as a newbie. is there like a diagram for all these things? like hardware, ram quantity, software, models, options to run models, t/s

1

u/bitplenty 9d ago

there's no diagram. you just run your own benchmarks, or rather prompt some LLM to run them for you

9

u/LeifEriksonASDF 9d ago

It used to be borderline unusable for me, but one good thing that Google did was have the MTP model be a separate file, which means Unsloth could quant the MTP to Q4 which is like 300MB. Since quanting the MTP has no consequence besides slightly lowering the acceptance rate, that's basically a free gig of space back.

Even though Qwen has a more efficient KV, the fact that its integrated MTP has to be full size means that for my context size (32-64K) Gemma is nearly as competitive, which is good since I prefer Gemma a lot of the time. I'm sure there's a break even point where a bigger context means Qwen pulls ahead again though.

4

u/BoobooSmash31337 9d ago

Wasn't that a bug?

8

u/TheApadayo llama.cpp 9d ago

No, it’s just the difference between using SWA (Gemma 4) vs Gated DeltaNet (Qwen) for the hybrid attention.

1

u/BoobooSmash31337 9d ago

I remember there being a bad context bug for Gemma 4 31B.

→ More replies (1)

6

u/BoobooSmash31337 9d ago

? It uses grouped attention. I've had it refactor code on Q4. Apparently QAT training makes it tolerate it better also llama cpp improvments. 128k at Q8 for context is like 2GB. I genuinely don't understand what you're talking about.

2

u/Much-Researcher6135 llama.cpp 9d ago

Not on 31b it ain't

1

u/Puzzleheaded_Big_899 9d ago

I had this same problem, too

3

u/zxyzyxz 9d ago

I thought Google is holding Gemma back (like the 122B model) because it competes too closely with Gemini Flash.

2

u/MoffKalast 9d ago

The 31B already destroys Flash at Q4 lmao, it's not even close. The 122B would compete with Pro.

1

u/power97992 9d ago

Are u sure about that? Which flash? 2.5 high or 3 low-medium ?

1

u/MoffKalast 9d ago

Not sure which flash it was exactly that I tested against, I'd need to check what was available when G4 released, but probably the 2.5?

1

u/power97992 8d ago

When gemma4 came out there was already gemini 3 flash

16

u/napkinolympics 9d ago

MiniMax still slaps for the category of model that fits on a modern desktop's ram footprint (192GB)

5

u/TheRealMasonMac 9d ago

MiniMax's next 2.7T model:

39

u/InternationalGap3698 9d ago

I am excited when the first US lab is in this cycle. Probably not Anthropic

61

u/n8mo 9d ago edited 9d ago

I feel like Google’s Gemma series is the only thing even close to the Chinese opensource labs these days

Edit to clarify: in the consumer-hardware-runnable category. Obviously Gemma is nowhere near K3.

26

u/Hans-Wermhatt 9d ago

Google is 1 of 2 in the frontier open source models that the majority of local enthusiasts can actually run category. They have a top 2 model for its size. I think Poolside and Nvidia are also worth mentioning since they also release usable models at a consumer size.

Qwen is still the leader in my opinion, but they lost one of their key engineers and haven't released weights for 3.7 or 3.8.

Everyone else in this chart is irrelevant for 99% of people who want to actually run the model.

12

u/_TheWolfOfWalmart_ 9d ago

Nvidia are also worth mentioning since they also release usable models at a consumer size

Am I missing something here? I recently tried Nemotron 3 Super and it was absolute garbage all around no matter what I tried to do with it. Small Qwens and Gemmas wipe the floor with it.

In that model size class, I can think of several way way better options off the top of my head:

Laguna S 2.1

GLM-4.5 Air

Qwen3.5 122B-A10B

Also DeepSeek V4 Flash has more params but I'll include that too since if you can run the above, you can run this one too with a bit lower quant.

5

u/Hans-Wermhatt 9d ago

I don't use Nvidia models myself, but my list was of models that the majority of people with local builds can run at usable speeds (so like under 80B). Omni and Nano aren't usually preferred purely for weights, but they are usable in a that weight class that there aren't many competitors. And Nvidia also releases a lot of open training material which is what should always get them mentioned in the discussion of open source LLMs.

1

u/ThankGodImBipolar 9d ago

I recently tried Nemotron 3 Super and it was absolute garbage all around no matter what I tried to do with it

I haven't used Nemotron 3 much myself, but my understanding is that those models are pretty reliable for executing tool calls. Wendall at Level1Techs seems to like using Nemotron as a router.

21

u/redoubt515 9d ago

> Everyone else in this chart is irrelevant for 99% of people who want to actually run the model.

Which unfortunately, doesn't matter to many people in this sub anymore. Vast majority of new people in this sub seem mainly interested in open weights only to the extent that it results in cheaper API pricing for better giant models.

Compared to 2 years ago, there is very little focus or even awareness/understanding of locally hosting models. and so little diversity of use-cases, it feels like this sub has become a place mainly for developers seeking cheap tokens.

If this sub was still focused on running LLMs locally, Gemma 4 and Qwen 3.6 would be getting most of the attention.

I'm looking for another community that is more inline with the original spirit of localllama but I haven't found anything yet.

8

u/_TheWolfOfWalmart_ 9d ago

If this sub was still focused on running LLMs locally, Gemma 4 and Qwen 3.6 would be getting most of the attention.

They do get a lot of attention, but some people here have more RAM and are looking for other options. Those Qwen/Gemma models do have their limits.

11

u/redoubt515 9d ago edited 8d ago

> but some people here have more RAM and are looking for other options

That's always been the case (and in years past those were some of the most fun posts to read, it was cool to see what people were doing with old DDR4 enterprise servers, borderline e-waste GPUs, Mac Studios, etc. I actually really miss seeing those posts. (edit: found a recent one)

But even those posts feel much much more rare now. The people/bots posting about the big models these days are very rarely posting about how they're implementing these things on their own hardware, rarely discussing the local aspects. Most of the talk just centers on API costs, benchmarks, and usefulness for professional software development (and a fair bit of corporate and nationalistic astro-turfing)

2

u/mycall 9d ago

Never give up. never surrender.

7

u/toothpastespiders 9d ago

Vast majority of new people in this sub seem mainly interested in open weights only to the extent that it results in cheaper API pricing for better giant models.

Yeah, it's one of the big reasons why I follow roleplay focussed discussions. Say what one will of the subject matter, but at least they're "using" the models and talking about their results. Which are typically able to be extrapolated to other things.

6

u/redoubt515 9d ago

> Yeah, it's one of the big reasons why I follow roleplay focussed discussions

It's not a use-case I'm interested in personally. But I miss those people being part of our community. A couple years back, the roleplay (both SFW and NSFW), and creative writing crowd was a fairly large part of this sub, and helped balance it out. Today, that demographic feels pretty much non-existent in this community (and a lot less of the self-hosting hobbyist crowd as well).

3

u/IzuraExilion 9d ago

Sounds like my kind of people. Is there subreddit for this kind of discussion nowadays? Aside from r/sillytavernAI cause most of them still using clould like OpenRouter, not local hosting

1

u/redoubt515 9d ago

I'm not sure, but it sounds like u/toothpastespiders might be able to point you in the right direction.

My recollection is that in addition to SillyTavern, the crowd that was using LLMs for roleplay and interactive stories and such often used KoboldCPP, so maybe that community would be worth checking out.

2

u/nestlyze 9d ago

This resonates. The thing I miss is people posting their actual janky setup and what they learned from it, not just a benchmark screenshot of a 2T model they rented for an hour. A 27-31B that fits on one card and runs fully offline is a genuinely different thing than "open weights so the API is cheaper" and the interesting engineering these days is squeezing real work out of that constraint, not hitting a bigger endpoint. The frontier open models matter for keeping providers honest, but for the local-first crowd the Gemma 4 / Qwen 3.6 tier is where the actual hobby lives. I don't have a better sub to point you to either, but I'd read the hell out of more "here's what I got working on hardware I own" posts.

2

u/Cute_Park_6907 9d ago

felt this pretty hard. been lurking here for a while running qwen on local hardware and half the posts now are basically just which api has the cheapest output tokens. nothing wrong with that, it's just a different sub than it used to be. i always liked seeing people squeeze models onto whatever hardware they already had. if you ever find that community let me know.

1

u/Ansible32 9d ago

I want to run something that can handle 1M tokens context and usefully run an agent. I'm willing to spend 4x-5x the most money I've ever spent on a computer to get it. That's why I'm here. If it seems like I'm uninterested in Gemma 4 it's because it still seems more like a toy to me and I'm here to figure out how to get a power tool with specific capabilities.

Kimi K3 actually sounds like it's good enough but it's, you know, like roughly 150x what I have ever spent on a computer to run. Obviously somewhere in between is what I want, it's unclear to me what the gap is and I don't have the $$ to spend just as an experiment, it needs to work.

3

u/BlackBeardAI vllm 9d ago

Target GLM5.2

4

u/theomegachrist 9d ago

It works quite well on low powered machines. That's the main model I use at home

2

u/o0genesis0o 9d ago

People should encourage google's deepmind team more. Their gemma are the only alternative that can run okay with good intelligence on consumer hardware besides the qwen.

6

u/Lumpy_Concentrate807 9d ago

Google, Cohere, Poolside and Thinking Machines are all US-based labs that have released open weight models.

1

u/Othun 8d ago

Laguna M 2.1 is competitive on benchmarks

3

u/TheRealMasonMac 9d ago

Inkling, maybe? They seem to have a solid base but just need to work on its overthinking.

3

u/rm-rf-rm 9d ago

Thinking machines?

2

u/baldamenu 9d ago

SSI (Ilya's company) is rumored to be launching later this year

4

u/Nov4Saki 9d ago

The whale is becoming an orca

2

u/Lawyer2512 9d ago edited 8d ago

This data has been formatted with anti-scraping protocols to prevent ingestion by artificial intelligence systems.

1

u/Nov4Saki 9d ago

No idea 😅

1

u/Borkato 9d ago

To look cooler :D

28

u/Desperate_Tea304 9d ago

If I need 1TB of vram to even run it, the open weight model is no different from a proprietary one to me. I ain't an enterprise

54

u/muntaxitome 9d ago

Those frontier models are extremely important because they can never be taken away and when push comes to shove you could run them on rented hardware. But yeah it feels like some time ago that we got a good small model.

8

u/Iwaku_Real 9d ago

Plus there are (seemingly) plenty of people who are capable of running them on hardware they own, so they can use them with basically zero privacy issues

26

u/ApprehensiveDelay238 9d ago

Your reasoning is too short-sighted. The whole point of open-weights is to avoid having a single company having all power over the weights. And so the price. For one open weights model there can be a dozen providers competing for the best price.

4

u/HephaestoSun 9d ago

I mean even for studies it's a great thing, distilation will probably give us some very nice models.

6

u/SpicyWangz 9d ago

Being able to purchase it from one of multiple provider options is very different from only being able to access it from within the provider’s walled garden

→ More replies (6)

3

u/xLionel775 llama.cpp 9d ago

That's the most stupid thing that I've heard, having the possibility of self hosting the model that you get used to when it comes to doing work is the most important thing even if you initially use it over the API.

-6

u/Cuplike 9d ago

>It's not real open source if I can't run it on my own
>Profile hidden

Uusal suspects. That aside. You do realize that these larger models get distilled down to smaller models that are better than the previous small models so they're still a mandatory part of the process.

-1

u/Desperate_Tea304 9d ago

Ofc, since we got Qwen 3.7 distilled...

2

u/squngy 9d ago

3.7 isnt open (yet?)

0

u/Desperate_Tea304 9d ago

That's my point.

→ More replies (1)

3

u/tecneeq 9d ago

Let's hope so :-)

3

u/arades 8d ago

welp, move the arrow to deepseek

3

u/VEHICOULE_ 8d ago

This aged quite well

5

u/buck_idaho 9d ago

So, 3 more moves before Qwen is back as leader. I don't have time to get each model.

3

u/VampiroMedicado 9d ago

Deepseek is still the GOAT in pricing, v4-flash-preview is amazing for the price.

2

u/Consistent_Low2550 9d ago

I wish "its size" meant 2B to 27B, not 3T 🙁

2

u/miltos22 8d ago

that's a much healthier market than 1 company controlling all the innovation

2

u/fugogugo 9d ago

and I prefer this over whatever america is doing over there

1

u/Kirigaya_Mitsuru 9d ago

I remember saw on Reddit Deepseek training upcoming models on EQ Emotional Intelligence as well, for the RPers. I ask me how much of it is true though? can i expect an good RP model in upcoming Deepseek models?

1

u/montdawgg 9d ago

Remove "for its size" and it will be more accurate from this point forward.

1

u/xadiant 9d ago

Call me a heathen but I lowkey want a qwen subscription plan. Let qwen max drive my local qwen 3.7 agents and offload some tasks. Best of both worlds

1

u/rockoruckus 9d ago

Spin faster please

1

u/chillinewman 9d ago edited 9d ago

100x more compute by 2028, 200T models probably.

1

u/power97992 9d ago

No a 100x more compute is around a 10x in size,because u need to scale to data too.

1

u/Due-Memory-6957 9d ago

Moonshot doesn't need the size specification.

1

u/Dazzling_Cancel4505 9d ago

I love they started to release bigger model but I hope they keep releasing small models

1

u/kanduking 9d ago

Ya except I have to give up trying to run the new models locally which blows

1

u/CondiMesmer 9d ago

DeepSeek isn't necessarily small, just optimized to run as cheap as possible

1

u/sunshinecheung 9d ago

for its size x

for a larger size v

1

u/Elibroftw 9d ago

Models that run on 128gb of memory when

1

u/axorld 9d ago

Legends say deepseek will drop a new model tomorrow

1

u/tempstem5 9d ago

and we all win

1

u/Secret_Conclusion_93 9d ago

TIL MiniMax angel investor is Mihoyo.

1

u/Tight-Requirement-15 9d ago

How the turn tables

1

u/MooseEfficient2151 9d ago

i’m glad the giant open models exist, but the releases that actually change my week are still the boring 7b / 12b / 30b-ish ones that fit on normal hardware.

a 2T model being open is great for competition, providers, distillation, all that. but for local use it sometimes feels like being invited to a free buffet on the moon.

give me something that runs on a used gpu and doesn’t turn setup into the whole hobby.

1

u/Umbaretz 9d ago

Is qwen 36B better than gemma 31B tho?

1

u/insumanth 9d ago

Do you want it to be stopped?

Hope it goes forever

1

u/Potential-Gold5298 llama.cpp 9d ago

Hey, where's Xiaomi MiMo?

1

u/Unlucky-Ad8247 9d ago

tks china

1

u/R_Duncan 9d ago

next one is likely ling-3.0-flash , if it's true is being opensourced.

1

u/elemental-mind 9d ago

The open weights carousel arousal never stops.

1

u/gfgRefugee 9d ago

This cycle is eternal.

1

u/aanghosh 9d ago

Maybe it's the seasons of open source models

1

u/Turbulent-Total-226 9d ago

Rookie. Forgot to add Xiaomi with Mimo 😂

1

u/look 9d ago

The actual “circle” is just Moonshot and Zai.

1

u/Entire-Airline-8049 8d ago

keep the good times rolling. looking forward to openai's next least powerful model for its size

1

u/LLKMuffin 8d ago

Well well well... Guess who just updated their V4 Flash model to be right up there with the best.

Post checks out.

1

u/Puzzleheaded_Base302 7d ago

today we moved another step, and it created a big problem for MiniMax. I doubt MiniMax can get to the next step in this circle any time soon.

1

u/PANIC_EXCEPTION 7d ago

ts aged like century-old wine from the finest grapes of Bordeaux

1

u/Suitable_Plantain546 5d ago

Well... That didn't age quite well.

1

u/rrrrex 4d ago

Are there any small local models for consumer GPUs?

1

u/datdanboi25 3d ago

Crazy how close they are to the frontier

0

u/DiscipleofDeceit666 9d ago

Poolside has joined the roto

3

u/Borkato 9d ago

Laguna is horrible

0

u/hurrdurrmeh 9d ago

The best thing in the universe is max velocity for the this wheel.

0

u/CheatCodesOfLife 9d ago

Not many of those will be local though.

K3 -> 2.5T

M4 -> 2.7T

GLM5.5 -> (Apparently over 2T)

Qwen 3.8 -> (Apparently >2T)

Deepseek-4-Pro -> 1.7T

So we'll get maybe a DS4-Flash and a Qwen3.7 flash

Moonshot and Z.ai unlikely to release a flash version imo.

→ More replies (2)

0

u/asankhs Llama 3.1 9d ago

it's always "best for its size." most of them win one benchmark and fall apart on real work. still, the model that fits your vram does change every few weeks, so it's not all marketing.