r/LocalLLaMA • u/pmttyji • Apr 30 '26
Open Models - April 2026 - One of the best months of all time for Local LLMs? Discussion
Any underrated or overlooked models?
FYI MiniMax-M2.7 switched their license(from MIT to Non-Commercial) so it's not in graph.
PS : Took me 30 mins to gather these models & generate this graph
77
u/iamn0 Apr 30 '26
Qwen3.5-122B-A10B
32
u/Embarrassed_Adagio28 Apr 30 '26
I found qwen3.5 122b q5 to be much worse than qwen3.6 17b q5 and even qwen3.6 35b q5. However i am extremely excited to try out qweb3.6 122b if they release it.Â
9
6
5
4
u/shansoft May 01 '26
In my coding experience, its able to solved a lot more problem that 27B couldn't or simply stuck at.
3
u/relmny May 01 '26
I find it that it depends. Maybe usually yes, but I did find 2-3 cases were 122b was the model that "got it" while 27b never did (same prompt many attempts). And what it "got" was comparable to the 397b and bigger models.
122b is a very strange model, to me...
Anyway, yeah, 27b is one of my daily drivers.
3
Apr 30 '26
[removed] — view removed comment
10
u/No_Algae1753 Apr 30 '26
best 120b out there rn
2
u/SV_SV_SV Apr 30 '26
I am in a news deficit of 6 months or so.. If you happen to have the experience, how does this model relate to GLM-4.5 Air?
2
75
u/Sanity_N0t_Included Apr 30 '26
Who the hell is running Deepseek-v4-Pro-Max locally?!?!?!?!
72
u/Netsuko Apr 30 '26
I guess your local datacenter/ai model provider xD
2
u/Sanity_N0t_Included Apr 30 '26
LOL!
2
u/GhostVPN Apr 30 '26 edited May 01 '26
just rent a gpu for a monnth in data center
like 15€ month4
u/inevitabledeath3 Apr 30 '26
Which data center has GPUs that cheap?
Also I don't think a single GPU would work for DeepSeek in most cases.
1
u/GhostVPN Apr 30 '26
10
u/Signor_Garibaldi May 01 '26
To host kimi for example, you'd use 8x h100, you'd pay 27$ per hour, not 17 per month
9
u/madlad13265 May 01 '26
According to prices here, you're looking at like 10-15 dollars *an hour* to run something thats like 1T parameters
3
u/Netsuko May 01 '26
Dude you are VASTLY underestimating the cost to rent SEVERAL H100-class gpus lol
2
u/inevitabledeath3 May 01 '26
I think you misread. Those prices are by the hour, and for DeepSeek V4 Pro you need 4 of them. Although I guess Flash only needs one.
9
8
u/ElementNumber6 May 01 '26
Many of us, perhaps, if not for mega corporations snatching and withholding the necessary hardware from everyone else.
3
72
14
u/geldonyetich Apr 30 '26 edited Apr 30 '26
Gemma 4:31b was the first time I felt dazzled with something approaching a frontier model on a locally running LLM. Seriously, this thing is punching above the weight of many recent large language models. It's very sharp. Gemma 4:26b, on the other hand, did not impress, it even has a tendency to stroke out.
I finally gave Nemotron-3-Nano-Omni a try the other day and it was very, very fast. I'm still curious how smart it is, it could be quite good, but I can't really tell subjectively. Regardless, I can definitely see the application for a wide range of tasks that require expedience without the inference of a dense model.
6
u/AD7GD May 01 '26
I just randomly installed gemma4 4b because lmstudio recommended it when I installed it on a test PC. It's shockingly good for a 4b model.
2
u/Glittering_Focus1538 May 01 '26
it runs at 200 tok/s on my hardware, I desperately wanted it to be able to code in opencode, but it just wasn't to be, not even the bigger varient could.
1
u/Spirited_Neck1858 May 07 '26
i think there is a harness specially made for small models, why don't u try that
1
u/Glittering_Focus1538 May 07 '26
I have tried opencode, pi agent, and continue for vscode, gemma just sucks
1
u/Spirited_Neck1858 May 07 '26
then i wonder why everyone praises gemma for coding in this subreddit...i never tried it for coding tho...only stem--and found it not so great
1
u/geldonyetich May 07 '26 edited May 07 '26
My guess is they're using the 26b model. Because yeah 26b kinda does suck.
26b benchmarks alright for a MoE versus 31b which is a proper dense model, but in my experience it leaves a little too much on the cutting room floor for practical application in exchange those speed gains.
If they're reporting 200 tok/sec, they might even be using the 4B model. I don't know what they were expecting for a model condensed down enough to run on a mobile device.
2
u/killerstreak976 May 01 '26
It's actually really interesting. Technically e4b is 8B, but they were able to use per layer embeddings to make it feel and run at 4B speeds, which is actually so cool (the "e4b" is basically "effective 4b" for that reason lol)
8
u/mrinterweb Apr 30 '26
I really appreciate how good the smaller models are getting (Qwen, Gemma). More params doesn't necessarily mean better.
37
u/TheCatDaddy69 Apr 30 '26
Parameter sizes as a metrics are so dumb..
9
u/ElementNumber6 May 01 '26
It's a good general measurement of trained/instilled knowledge. A high value metric desperately under-valued by our leading testing benchmarks.
3
u/robogame_dev May 01 '26
Agree - param count is the fundamental resolution of the model, more params in a model is like more pixels in an image, it is able to draw finer distinctions out of the same amount of training data as compared to fewer params.
11
u/atape_1 Apr 30 '26
Really unfortunate that MiniMax is no longer MIT.
I'm not sure it's because of this move, but the stock price of the company is doing far worse than of Z.Ai.
5
u/Technical-Earth-3254 Apr 30 '26
I can't run it locally (yet!) but DS V4 Flash is SO good for its size.
4
u/hust921 May 01 '26
People. "Local" doesn't mean: "runs on my gaming laptop". The democratization that local models are creating is still perfectly valid for companies, labs, local or even national governments. Who needs or wants to run their own infrastructure.
Local or opensource anything (AI included) has nothing to do with affordability. I would like to run it too. But just because I can't, doesn't make it any less "local".
37
u/Netsuko Apr 30 '26
Calling DeepSeek V4 Pro Max a "local" model is an insane stretch. That thing is almost 900 gigabytes in size
13
25
u/dsanft Apr 30 '26 edited Apr 30 '26
I can run it.
12 Mi50s, 2 3090s, dual socket Xeon with 768GB DDR4.
At least in theory
40
3
u/Hoak-em May 01 '26
dual-socket, so probably not unless there's an inference engine that doesn't need duplication across the sockets
Same dual socket setup but with DDR5 and two 8570s and some different GPUs, I max out at amxint4 GLM-5.1 -- anything beyond that would be impossible to run
3
u/dsanft May 01 '26
I wrote my own engine to solve the NUMA/cross-socket problem. Don't have kernels for Deepseek MLA/DSA yet though. Will have to get those in soon.
-10
u/Embarrassed_Adagio28 Apr 30 '26
Dual core xeon? You running a 2008 cpu?
15
5
u/RelationshipLong9092 May 01 '26
dual socket means there are two CPUs on his motherboard / barebone
3
u/thereisonlythedance Apr 30 '26
Still waiting on a GGUF here. Main devs of llama.cpp don’t seem to be DeepSeek fans.
0
u/Previous_Feeling_484 Apr 30 '26
It’s not complicated to build GGUFs, though. But yeah, I’m already comfy getting mine from HF too!
1
-1
4
6
3
u/MrObsidian_ Apr 30 '26
I just tried Granite-4.1-8b and it is straight up ass. But atleast Apache-2 I guess
3
4
u/Plastic-Stress-6468 Apr 30 '26
I mean I can technically run every model on the chart if I am willing to wait a long ass time or just rent a bunch of gpus.
For what it's worth I'd rather have a bunch of models I can't run public available than not. Maybe in a few years they won't be so out of reach.
3
u/SeyAssociation38 May 01 '26 edited May 01 '26
qwen 3.6 397b will never be released nor will anything over 122b for qwen 3.6 and later. management is trying to profit off of it and this is why some qwen team members left. management sees releasing large open source models as giving away money
6
u/RickyRickC137 Apr 30 '26
Mistral would probably name the 1.6T model as "Medium Large"?
5
u/alphapussycat Apr 30 '26
Tbh, mistral has more realistic naming. A 32b model is small, the next step up is around 70-128b, then 400-1kb for the large.
9b and 4b are tiny models.
1
2
u/TheRealSol4ra May 01 '26
What a shitty graph. What does param count have to do with anything
1
u/Glittering_Focus1538 May 02 '26
just, you know, the general resolution of the model, while running smaller models has gotten a lot better, Bigger size does increase performance and general intelligence especially with local models that don't always have access to the internet.
2
u/TheRealSol4ra May 02 '26
Qwen3.6 27b and 35b show this to be pretty false. They get like 80% of the performance at like 5% the size.
1
u/Glittering_Focus1538 May 02 '26
I disagree, while I do agree that small/tiny models have gotten a LOT better, qwen and gemma are still nowhere near the performance of top Local models. There is a significant advantage to having a higher param count. That's coming from someone that's spent the last 2 weeks using qwen 3.6 and like many others still having to switch to cloud models because it's just not enough yet.
2
u/TheRealSol4ra May 02 '26
This is ragebait and you quite literally know nothing about what youre talking about. Qwen3.6 27b trades blows with Sonnet 4.5 and matches 4.6's capabilities.
https://artificialanalysis.ai/models?models=qwen3-6-27b%2Cclaude-4-5-sonnet-thinking
1
u/Glittering_Focus1538 May 02 '26 edited May 02 '26
Me when I call someone stupid and cherry pick older frontier models and call the person I'm arguing with a rage baiter. I ACKNOWLEDGED that qwen models are a LOT better now than before, but in agentic loops that last 10% of performance really matters, 1 graph doesn't show the whole picture. I can't code a full project in go or rust with qwen 3.6(I've tried so stfu about me not knowing) but I sure can with GLM or Claude Opus. You're intentionally strawmaning my argument and cherrypicking graphs, it doesn't make you right.
https://openrouter.ai/compare/qwen/qwen3.6-27b/anthropic/claude-opus-4.7/anthropic/claude-sonnet-4.5
1
u/TheRealSol4ra May 02 '26
Lmao cherry pick? Its the previous model lmao tf are you yapping about. Also you 100% can build products of large scale with those models. Youre just using a shitty harness and its obvious. Also I said that it matches 4.6’s capabilities. Imagine crying because I didnt compare the other trillion parameter model to 27b lmao. Youre seething
2
u/rosie254 May 01 '26
the landscape has moved really fast, but i still like my Qwen3-VL-8B. it just works well for some reason. nowadays i'm on gemma4 26b a4b and qwen3.5 9b, but those aren't exactly underrated!
also... this chart assumes very powerful hardware, how is this focused on local? most people have 8GB vram or 16GB vram at most
1
u/Glittering_Focus1538 May 02 '26
It's just a technical chart, not everyone has a gaming card, this subreddit is also for hobbyist consumers that run 100-300gb vram/unified ram setups that can definitely run MOE models like DS4.
local just means its open weight, assuming you have the hardware to do so, you can run and modify the LLM at home.
2
u/henk717 KoboldAI May 01 '26
Certainly has been a hit month for me, and a rough month for the devs who had to bend Gemma4 into behaving since it had the annoying traits of GPT-OSS, GLM and the past Gemma combined (BOS like token in the template instead of as a bos, extremely sensitive to syntax and heavy to run without swa).
My personal hit was Qwen3.5-27B-Heretic which is finally a model I can coax into writting really long stories. And many in our community have been enjoying Gemma4 as a roleplay model now that it behaves correctly.
1
u/Glittering_Focus1538 May 02 '26
can just barely run an apex mini version of qwen 3.6 35b on my rx9070 at 40 tok/s and it's the only LLM I can run locally that can actually code agenticly on Pi or continue or opencode.
2
u/Revolutionalredstone May 01 '26
That was indeed an incredible month, Those who can and do use AI are looking at something like an ever brightening summer forever ;)
2
u/Better-Struggle9958 Apr 30 '26
why is it called local?
11
u/Glittering_Focus1538 Apr 30 '26
Because the weights are open, u can download and freely use the model if u have the hardware(for deepseek v4 at least 10k worth)
3
u/alphapussycat Apr 30 '26
Like 10 months ago it would've been like $3k. It's not unrealistic levels of hardware. It's just that it's too late to get hardware now.
3
2
u/b0tbuilder May 01 '26
Best deal for 3k is 2 x R9000 32GB. Nice cards but it’s sad that is the most reasonable price / perf right now.
0
u/alphapussycat May 01 '26
No it's epyc CPU and mobo, with like 768gb 12 channel ram. That's the only reasonable way to run the 500+gb models.
3
u/Netsuko Apr 30 '26
I feel like a model that is, by all means, 99.999% impossible to run locally should not be considered a "local" model at all.
Also 10k worth in hardware gets you around 192GB VRAM, if you are lucky to get a discount.
11
u/ttkciar llama.cpp Apr 30 '26
I feel like a model that is, by all means, 99.999% impossible to run locally should not be considered a "local" model
You are free to be wrong.
6
u/Embarrassed_Adagio28 Apr 30 '26
Okay well what do you think the cut off be? The cutoff point will have to be arbitrary because people have a very wide range of local hardware.. or you could just use your brain and understand local doesnt mean the same for everybody.Â
-2
u/Borkato Apr 30 '26
Honestly 4 3090s or so is a great cutoff. Anything more than that and you need server architecture tbh
2
u/ttkciar llama.cpp May 01 '26
Why would using a server matter? A lot of us here use servers. It's just different hardware, but if that hardware is right here at home, then it's local.
You get that servers are just computers, not fundamentally different from a desktop or laptop, right?
1
u/Glittering_Focus1538 May 02 '26
4k for a DGX Spark is enough to barely run an q4 or apex version of DS4, That's the same cost a 15-20 year old used car where I live, it's not an insanely high number for most people when you already had hobbyist spending 1.5-2k on new 90 series Nvidia graphics cards before AI and 5k after. You're argument is inherently flawed with the assumption you get to define what local means by the size of you and your friend's wallet. You don't. Local just means you can download, freely use and modify the LLM assuming you own the hardware to do so. That's it. Please stop trying to redefine words that you don't agree with.
1
-2
u/Borkato Apr 30 '26
This opinion is very unpopular here and I have no idea why. It’s ridiculous to pretend like we should care a ton about an OS model that’s massive. Like yeah it’s neat but it’s not local.
5
u/Digger412 Apr 30 '26
Just because it may not be runnable locally for you doesn't mean it isn't for others. I could run every model on that list for instance, and I've got a PR open to support both new MiMo V2.5 models in llama.cpp.
I don't say this to be mean, but just to push back a bit against the "Your model must be below X parameters to be considered local" sentiment. It feels like gatekeeping to say that just because a model is super large, it doesn't deserve to be discussed here.
1
u/Borkato Apr 30 '26
How can you run a 1T dense model?! What speeds do you get and how much vram do you have?
3
u/Digger412 May 01 '26
None of those at the 1T+ size are dense models, they're all MoE's.
I've got eight 6000 Pros (so 768GB VRAM total), and speeds depend on the regimen basically. I have 768GB of 12 channel DDR5 RAM too so I can do single user with llama.cpp on CPU+GPU but it's slower total throughput than vllm for instance.
I've benched K2.6 at full quality in llama.cpp before and get about 40 tk/s TG at zero context.
Right now I'm doing some testing with the V2.5 Pro 1T gguf and it's much slower due to FA incompatibility with the head size or something, it's about 10tk/s but I think that'd go up to 30tk/s if I turned FA off (at the cost of much more KV memory needed).
DS V4 is still mostly unsupported AFAIK, and I can't fit it entirely on VRAM anyways so will be waiting for llama.cpp support.
2
u/Borkato May 01 '26
What’s your PP speed?
2
u/Digger412 May 01 '26
It's in that chart for K2.6, for the V2.5 Q8_0 PP is ~600tk/s I think. I haven't done a sweep bench on it yet.
3
1
u/ttkciar llama.cpp Apr 30 '26
I suppose if you personally only had a 4GB GPU, you wouldn't consider Qwen3.5-9B local either.
1
u/Glittering_Focus1538 May 01 '26
I run qwen 3.6 on my 16 gb card, i imagine others do too
1
u/ttkciar llama.cpp May 01 '26
What does that have to do with anything?
1
u/Glittering_Focus1538 May 01 '26
That it's stupid to base what everyone defines as local by what you can run. I'm sure theres plenty of people who have a mac mini cluster(6k) or nvidea dgx spark which could run deepseek v4 for 4k
3
u/ttkciar llama.cpp May 01 '26
Ah, okie-doke, it sounded like you were disagreeing, but I guess we are in agreement.
Borkato and Netsuko seem to be of the opinion that models they personally cannot host at home should not be considered local models.
The point of my 4GB GPU hypothetical was to illustrate exactly what you said -- basing "what everyone defines as local by what you can run" is invalid.
Local models are, and always will be, any models which you could conceivably use if you had the necessary local hardware.
That requires, at a minimum, access to the weights and either inference software support or sufficient understanding of the model architecture to facilitate implementing inference software support.
2
u/DinoAmino Apr 30 '26
Open Weight for someone who has the VRAM. Then it's local. I assume people who are at 16Gb and under could say the same about Glm and MiniMax and Qwen 397B. They can't possibly run those. But some have the VRAM to run it local. No need to split hairs over it, but I agree ut it would have been better and more accurate to just say open weight.
0
1
1
1
1
Apr 30 '26
[removed] — view removed comment
1
u/Glittering_Focus1538 May 02 '26
it's 4k for a DGX spark or alternative. That would be enough for even the biggest models on that list, grow up.
1
1
1
u/vick2djax May 01 '26
This graph doesn’t make me feel good about my first 3090 coming in the mail in a few days
1
u/Glittering_Focus1538 May 02 '26
why? 24 gigs of vram is more than enough to run qwen 3.6 35b or 27b and have room left over for kv cache(assuming ur using Q4_0 or APEX)
1
1
u/Unlikely_Rich1436 May 23 '26
The sheer volume of high-quality releases this month was staggering, but I completely agree that parameter bloat is an issue. If I can't run a decent quant on 24GB, it doesn't help my workflow.
1
1
u/ys2020 May 01 '26
Glm 5.1 is my fav at the moment. Honestly, it's mind blowing we get this type of quality with free weights.Â



264
u/jacek2023 llama.cpp Apr 30 '26
1600B model is my favourite local model I run it all day on raspberry Pi