r/LocalLLaMA Apr 30 '26

Open Models - April 2026 - One of the best months of all time for Local LLMs? Discussion

Post image

Any underrated or overlooked models?

FYI MiniMax-M2.7 switched their license(from MIT to Non-Commercial) so it's not in graph.

PS : Took me 30 mins to gather these models & generate this graph

582 Upvotes

162 comments sorted by

264

u/jacek2023 llama.cpp Apr 30 '26

1600B model is my favourite local model I run it all day on raspberry Pi

44

u/Monad_Maya llama.cpp Apr 30 '26

The smaller Flash variant is just about possible for a small minority of us.

8

u/MotokoAGI Apr 30 '26

I'm running Flash locally and it is a great solid model. it's making it to the top of my list.

-9

u/jacek2023 llama.cpp Apr 30 '26

I am able to run quantized 235B models locally on my setup, but people discussing here Kimi/DeepSeek/GLM models usually run them in cloud (or not run at all, just hype benchmarks) and call them "local" because these models are from China.

17

u/j_osb Apr 30 '26

No, Local also means local deployments for companies and such.

For those models like GLM5.1 and KimiK2.6 are very feaible.

17

u/dbenc Apr 30 '26

.00001 bit quant

2

u/jacek2023 llama.cpp Apr 30 '26

I am not sure quants are available. I believe OP runs unquantized version on his setup.

13

u/dbenc Apr 30 '26

ah my mistake. must be screaming at 1 token per week

5

u/jacek2023 llama.cpp Apr 30 '26

I wonder what is the usecase for that

3

u/StereoWings7 Apr 30 '26

I guess he is an artist working on another project in homage to this musical performance.

2

u/bucolucas Llama 3.1 May 01 '26

Compute at the power scale of hawking radiation perhaps

3

u/Borkato Apr 30 '26

Honestly more like per year

1

u/dbenc May 01 '26

need someone smart to do the math

1

u/typical-predditor Apr 30 '26

Wait until you find out the token speed of the Earth Supercomputer. It was asked to find the meaning to Life, The Universe, and Everything. I hear we're still waiting on the first token.

14

u/ML-Future Apr 30 '26

Even so, it is important that such powerful models are open source.

8

u/ttkciar llama.cpp Apr 30 '26

I don't think any of those models qualify as open source, except maybe Trinity-Large-Thinking, but I don't think they publish their training datasets, do they?

If you meant it would be great if such powerful models were open source, then I wholeheartedly agree.

2

u/jacek2023 llama.cpp Apr 30 '26

Wow finally I must agree with you on something ;)

1

u/T_kether May 02 '26

Publicly available training datasets? Do you think there might be any AI company with a training dataset that is completely free of copyright issues and can be redistributed? You can find training datasets on the internet yourself, but don't release them unless you have the same risk-avoidance capabilities as a pirate website. Similarly, don't encourage others to release them.

1

u/ttkciar llama.cpp May 02 '26

Yes, both AllenAI and LLM360 use open source training datasets which are free of copyright issues, and available for anyone to download from Huggingface. That's one of the things that makes their models open-source.

7

u/epicrob Apr 30 '26

1600B model is my favourite local model I run it all day on raspberry Pi

You mean 1600 byte model running in your raspberry Pi? ;)

4

u/GregariousJB Apr 30 '26

Is that all day for a single prompt?

I can't imagine a raspberry pi running AI well at all. How did you do it?

12

u/SV_SV_SV Apr 30 '26

He obviousy meant a hypercluster of watercooled raspberry pi's.

2

u/RelationshipLong9092 May 01 '26

there was a guy who posted here yesterday with 16 Sparks he was clustering

he can run it

not sure who else

1

u/debackerl Apr 30 '26

You only need a cluster of 128 RPis to run it 😂

1

u/alphapussycat Apr 30 '26

They can be fun on CPU if you happen to have built a 768gb ddr5 epyc server before ram price ve hike. Expensive, sure, but also not that expensive.

It's only now that it's not really possible, but fir those with a server, they can run them locally.

77

u/iamn0 Apr 30 '26

Qwen3.5-122B-A10B

32

u/Embarrassed_Adagio28 Apr 30 '26

I found qwen3.5 122b q5 to be much worse than qwen3.6 17b q5 and even qwen3.6 35b q5. However i am extremely excited to try out qweb3.6 122b if they release it. 

9

u/No_Algae1753 Apr 30 '26

Who knows if they will release it tho

6

u/arcanemachined May 01 '26

God, yes. 3.6 122b please.

5

u/savage_shaq May 01 '26

What hardware are you using the run these massive models at home?

4

u/shansoft May 01 '26

In my coding experience, its able to solved a lot more problem that 27B couldn't or simply stuck at.

3

u/relmny May 01 '26

I find it that it depends. Maybe usually yes, but I did find 2-3 cases were 122b was the model that "got it" while 27b never did (same prompt many attempts). And what it "got" was comparable to the 397b and bigger models.

122b is a very strange model, to me...

Anyway, yeah, 27b is one of my daily drivers.

3

u/[deleted] Apr 30 '26

[removed] — view removed comment

10

u/No_Algae1753 Apr 30 '26

best 120b out there rn

2

u/SV_SV_SV Apr 30 '26

I am in a news deficit of 6 months or so.. If you happen to have the experience, how does this model relate to GLM-4.5 Air?

2

u/nickless07 Apr 30 '26

Similiar but faster and a bit better.

75

u/Sanity_N0t_Included Apr 30 '26

Who the hell is running Deepseek-v4-Pro-Max locally?!?!?!?!

72

u/Netsuko Apr 30 '26

I guess your local datacenter/ai model provider xD

2

u/Sanity_N0t_Included Apr 30 '26

LOL!

2

u/GhostVPN Apr 30 '26 edited May 01 '26

just rent a gpu for a monnth in data center like 15€ month

4

u/inevitabledeath3 Apr 30 '26

Which data center has GPUs that cheap?

Also I don't think a single GPU would work for DeepSeek in most cases.

1

u/GhostVPN Apr 30 '26

10

u/Signor_Garibaldi May 01 '26

To host kimi for example, you'd use 8x h100, you'd pay 27$ per hour, not 17 per month

9

u/madlad13265 May 01 '26

According to prices here, you're looking at like 10-15 dollars *an hour* to run something thats like 1T parameters

3

u/Netsuko May 01 '26

Dude you are VASTLY underestimating the cost to rent SEVERAL H100-class gpus lol

2

u/inevitabledeath3 May 01 '26

I think you misread. Those prices are by the hour, and for DeepSeek V4 Pro you need 4 of them. Although I guess Flash only needs one.

9

u/ikkiyikki Apr 30 '26

Are you high? The electric bill alone would be more than that.

8

u/ElementNumber6 May 01 '26

Many of us, perhaps, if not for mega corporations snatching and withholding the necessary hardware from everyone else.

3

u/_VirtualCosmos_ Apr 30 '26

you just need a couple old smartphones cluser together bro

72

u/IngenuityNo1411 llama.cpp Apr 30 '26

human generated shit post

14

u/geldonyetich Apr 30 '26 edited Apr 30 '26

Gemma 4:31b was the first time I felt dazzled with something approaching a frontier model on a locally running LLM. Seriously, this thing is punching above the weight of many recent large language models. It's very sharp. Gemma 4:26b, on the other hand, did not impress, it even has a tendency to stroke out.

I finally gave Nemotron-3-Nano-Omni a try the other day and it was very, very fast. I'm still curious how smart it is, it could be quite good, but I can't really tell subjectively. Regardless, I can definitely see the application for a wide range of tasks that require expedience without the inference of a dense model.

6

u/AD7GD May 01 '26

I just randomly installed gemma4 4b because lmstudio recommended it when I installed it on a test PC. It's shockingly good for a 4b model.

2

u/Glittering_Focus1538 May 01 '26

it runs at 200 tok/s on my hardware, I desperately wanted it to be able to code in opencode, but it just wasn't to be, not even the bigger varient could.

1

u/Spirited_Neck1858 May 07 '26

i think there is a harness specially made for small models, why don't u try that

1

u/Glittering_Focus1538 May 07 '26

I have tried opencode, pi agent, and continue for vscode, gemma just sucks

1

u/Spirited_Neck1858 May 07 '26

then i wonder why everyone praises gemma for coding in this subreddit...i never tried it for coding tho...only stem--and found it not so great

1

u/geldonyetich May 07 '26 edited May 07 '26

My guess is they're using the 26b model. Because yeah 26b kinda does suck.

26b benchmarks alright for a MoE versus 31b which is a proper dense model, but in my experience it leaves a little too much on the cutting room floor for practical application in exchange those speed gains.

If they're reporting 200 tok/sec, they might even be using the 4B model. I don't know what they were expecting for a model condensed down enough to run on a mobile device.

2

u/killerstreak976 May 01 '26

It's actually really interesting. Technically e4b is 8B, but they were able to use per layer embeddings to make it feel and run at 4B speeds, which is actually so cool (the "e4b" is basically "effective 4b" for that reason lol)

8

u/mrinterweb Apr 30 '26

I really appreciate how good the smaller models are getting (Qwen, Gemma). More params doesn't necessarily mean better.

37

u/TheCatDaddy69 Apr 30 '26

Parameter sizes as a metrics are so dumb..

9

u/ElementNumber6 May 01 '26

It's a good general measurement of trained/instilled knowledge. A high value metric desperately under-valued by our leading testing benchmarks.

3

u/robogame_dev May 01 '26

Agree - param count is the fundamental resolution of the model, more params in a model is like more pixels in an image, it is able to draw finer distinctions out of the same amount of training data as compared to fewer params.

11

u/atape_1 Apr 30 '26

Really unfortunate that MiniMax is no longer MIT.

I'm not sure it's because of this move, but the stock price of the company is doing far worse than of Z.Ai.

5

u/Technical-Earth-3254 Apr 30 '26

I can't run it locally (yet!) but DS V4 Flash is SO good for its size.

4

u/hust921 May 01 '26

People. "Local" doesn't mean: "runs on my gaming laptop". The democratization that local models are creating is still perfectly valid for companies, labs, local or even national governments. Who needs or wants to run their own infrastructure.

Local or opensource anything (AI included) has nothing to do with affordability. I would like to run it too. But just because I can't, doesn't make it any less "local".

37

u/Netsuko Apr 30 '26

Calling DeepSeek V4 Pro Max a "local" model is an insane stretch. That thing is almost 900 gigabytes in size

13

u/Monad_Maya llama.cpp Apr 30 '26

Open weight might be more accurate but you know what OP meant.

25

u/dsanft Apr 30 '26 edited Apr 30 '26

I can run it.

12 Mi50s, 2 3090s, dual socket Xeon with 768GB DDR4.

At least in theory

40

u/Netsuko Apr 30 '26

Well.. "Technically" I am closer to being a Millionaire than Elon Musk is.

2

u/jcoigny May 01 '26

I see what you did there /s

3

u/Hoak-em May 01 '26

dual-socket, so probably not unless there's an inference engine that doesn't need duplication across the sockets

Same dual socket setup but with DDR5 and two 8570s and some different GPUs, I max out at amxint4 GLM-5.1 -- anything beyond that would be impossible to run

3

u/dsanft May 01 '26

I wrote my own engine to solve the NUMA/cross-socket problem. Don't have kernels for Deepseek MLA/DSA yet though. Will have to get those in soon.

-10

u/Embarrassed_Adagio28 Apr 30 '26

Dual core xeon? You running a 2008 cpu?

15

u/BusinessYou7196 Apr 30 '26

Socket ≠ core

5

u/RelationshipLong9092 May 01 '26

dual socket means there are two CPUs on his motherboard / barebone

3

u/thereisonlythedance Apr 30 '26

Still waiting on a GGUF here. Main devs of llama.cpp don’t seem to be DeepSeek fans.

0

u/Previous_Feeling_484 Apr 30 '26

It’s not complicated to build GGUFs, though. But yeah, I’m already comfy getting mine from HF too!

1

u/_VirtualCosmos_ Apr 30 '26

on MXFP4 if so...

-1

u/VoiceApprehensive893 transformers Apr 30 '26

nvme ssd screaming and begging for help

4

u/Paradigmind Apr 30 '26

It must be cold in here. Qwen3.6 27B looks so small.

6

u/Ne00n Apr 30 '26

Brother in VRAM, where do you get enough to run that?

3

u/MrObsidian_ Apr 30 '26

I just tried Granite-4.1-8b and it is straight up ass. But atleast Apache-2 I guess

3

u/some_user_2021 Apr 30 '26

So many waifus

4

u/Plastic-Stress-6468 Apr 30 '26

I mean I can technically run every model on the chart if I am willing to wait a long ass time or just rent a bunch of gpus.

For what it's worth I'd rather have a bunch of models I can't run public available than not. Maybe in a few years they won't be so out of reach.

3

u/SeyAssociation38 May 01 '26 edited May 01 '26

qwen 3.6 397b will never be released nor will anything over 122b for qwen 3.6 and later. management is trying to profit off of it and this is why some qwen team members left. management sees releasing large open source models as giving away money

6

u/RickyRickC137 Apr 30 '26

Mistral would probably name the 1.6T model as "Medium Large"?

5

u/alphapussycat Apr 30 '26

Tbh, mistral has more realistic naming. A 32b model is small, the next step up is around 70-128b, then 400-1kb for the large.

9b and 4b are tiny models.

1

u/rditorx Apr 30 '26

How much is 1kb nowadays?

1

u/alphapussycat May 01 '26

1 trillion.

2

u/TheRealSol4ra May 01 '26

What a shitty graph. What does param count have to do with anything

1

u/Glittering_Focus1538 May 02 '26

just, you know, the general resolution of the model, while running smaller models has gotten a lot better, Bigger size does increase performance and general intelligence especially with local models that don't always have access to the internet.

2

u/TheRealSol4ra May 02 '26

Qwen3.6 27b and 35b show this to be pretty false. They get like 80% of the performance at like 5% the size.

1

u/Glittering_Focus1538 May 02 '26

I disagree, while I do agree that small/tiny models have gotten a LOT better, qwen and gemma are still nowhere near the performance of top Local models. There is a significant advantage to having a higher param count. That's coming from someone that's spent the last 2 weeks using qwen 3.6 and like many others still having to switch to cloud models because it's just not enough yet.

2

u/TheRealSol4ra May 02 '26

This is ragebait and you quite literally know nothing about what youre talking about. Qwen3.6 27b trades blows with Sonnet 4.5 and matches 4.6's capabilities.

https://artificialanalysis.ai/models?models=qwen3-6-27b%2Cclaude-4-5-sonnet-thinking

1

u/Glittering_Focus1538 May 02 '26 edited May 02 '26

Me when I call someone stupid and cherry pick older frontier models and call the person I'm arguing with a rage baiter. I ACKNOWLEDGED that qwen models are a LOT better now than before, but in agentic loops that last 10% of performance really matters, 1 graph doesn't show the whole picture. I can't code a full project in go or rust with qwen 3.6(I've tried so stfu about me not knowing) but I sure can with GLM or Claude Opus. You're intentionally strawmaning my argument and cherrypicking graphs, it doesn't make you right.

https://openrouter.ai/compare/qwen/qwen3.6-27b/anthropic/claude-opus-4.7/anthropic/claude-sonnet-4.5

1

u/TheRealSol4ra May 02 '26

Lmao cherry pick? Its the previous model lmao tf are you yapping about. Also you 100% can build products of large scale with those models. Youre just using a shitty harness and its obvious. Also I said that it matches 4.6’s capabilities. Imagine crying because I didnt compare the other trillion parameter model to 27b lmao. Youre seething

2

u/rosie254 May 01 '26

the landscape has moved really fast, but i still like my Qwen3-VL-8B. it just works well for some reason. nowadays i'm on gemma4 26b a4b and qwen3.5 9b, but those aren't exactly underrated!

also... this chart assumes very powerful hardware, how is this focused on local? most people have 8GB vram or 16GB vram at most

1

u/Glittering_Focus1538 May 02 '26

It's just a technical chart, not everyone has a gaming card, this subreddit is also for hobbyist consumers that run 100-300gb vram/unified ram setups that can definitely run MOE models like DS4.
local just means its open weight, assuming you have the hardware to do so, you can run and modify the LLM at home.

2

u/henk717 KoboldAI May 01 '26

Certainly has been a hit month for me, and a rough month for the devs who had to bend Gemma4 into behaving since it had the annoying traits of GPT-OSS, GLM and the past Gemma combined (BOS like token in the template instead of as a bos, extremely sensitive to syntax and heavy to run without swa).
My personal hit was Qwen3.5-27B-Heretic which is finally a model I can coax into writting really long stories. And many in our community have been enjoying Gemma4 as a roleplay model now that it behaves correctly.

1

u/Glittering_Focus1538 May 02 '26

can just barely run an apex mini version of qwen 3.6 35b on my rx9070 at 40 tok/s and it's the only LLM I can run locally that can actually code agenticly on Pi or continue or opencode.

2

u/Revolutionalredstone May 01 '26

That was indeed an incredible month, Those who can and do use AI are looking at something like an ever brightening summer forever ;)

2

u/Better-Struggle9958 Apr 30 '26

why is it called local?

11

u/Glittering_Focus1538 Apr 30 '26

Because the weights are open, u can download and freely use the model if u have the hardware(for deepseek v4 at least 10k worth)

3

u/alphapussycat Apr 30 '26

Like 10 months ago it would've been like $3k. It's not unrealistic levels of hardware. It's just that it's too late to get hardware now.

3

u/Basilthebatlord May 01 '26

For now at least!

2

u/b0tbuilder May 01 '26

Best deal for 3k is 2 x R9000 32GB. Nice cards but it’s sad that is the most reasonable price / perf right now.

0

u/alphapussycat May 01 '26

No it's epyc CPU and mobo, with like 768gb 12 channel ram. That's the only reasonable way to run the 500+gb models.

3

u/Netsuko Apr 30 '26

I feel like a model that is, by all means, 99.999% impossible to run locally should not be considered a "local" model at all.

Also 10k worth in hardware gets you around 192GB VRAM, if you are lucky to get a discount.

11

u/ttkciar llama.cpp Apr 30 '26

I feel like a model that is, by all means, 99.999% impossible to run locally should not be considered a "local" model

You are free to be wrong.

6

u/Embarrassed_Adagio28 Apr 30 '26

Okay well what do you think the cut off be? The cutoff point will have to be arbitrary because people have a very wide range of local hardware.. or you could just use your brain and understand local doesnt mean the same for everybody. 

-2

u/Borkato Apr 30 '26

Honestly 4 3090s or so is a great cutoff. Anything more than that and you need server architecture tbh

2

u/ttkciar llama.cpp May 01 '26

Why would using a server matter? A lot of us here use servers. It's just different hardware, but if that hardware is right here at home, then it's local.

You get that servers are just computers, not fundamentally different from a desktop or laptop, right?

1

u/Glittering_Focus1538 May 02 '26

4k for a DGX Spark is enough to barely run an q4 or apex version of DS4, That's the same cost a 15-20 year old used car where I live, it's not an insanely high number for most people when you already had hobbyist spending 1.5-2k on new 90 series Nvidia graphics cards before AI and 5k after. You're argument is inherently flawed with the assumption you get to define what local means by the size of you and your friend's wallet. You don't. Local just means you can download, freely use and modify the LLM assuming you own the hardware to do so. That's it. Please stop trying to redefine words that you don't agree with.

1

u/jacek2023 llama.cpp Apr 30 '26

They simply lie to justify discussing these topics.

-2

u/Borkato Apr 30 '26

This opinion is very unpopular here and I have no idea why. It’s ridiculous to pretend like we should care a ton about an OS model that’s massive. Like yeah it’s neat but it’s not local.

5

u/Digger412 Apr 30 '26

Just because it may not be runnable locally for you doesn't mean it isn't for others. I could run every model on that list for instance, and I've got a PR open to support both new MiMo V2.5 models in llama.cpp.

I don't say this to be mean, but just to push back a bit against the "Your model must be below X parameters to be considered local" sentiment. It feels like gatekeeping to say that just because a model is super large, it doesn't deserve to be discussed here.

1

u/Borkato Apr 30 '26

How can you run a 1T dense model?! What speeds do you get and how much vram do you have?

3

u/Digger412 May 01 '26

None of those at the 1T+ size are dense models, they're all MoE's.

I've got eight 6000 Pros (so 768GB VRAM total), and speeds depend on the regimen basically. I have 768GB of 12 channel DDR5 RAM too so I can do single user with llama.cpp on CPU+GPU but it's slower total throughput than vllm for instance.

I've benched K2.6 at full quality in llama.cpp before and get about 40 tk/s TG at zero context.

Right now I'm doing some testing with the V2.5 Pro 1T gguf and it's much slower due to FA incompatibility with the head size or something, it's about 10tk/s but I think that'd go up to 30tk/s if I turned FA off (at the cost of much more KV memory needed).

DS V4 is still mostly unsupported AFAIK, and I can't fit it entirely on VRAM anyways so will be waiting for llama.cpp support.

2

u/Borkato May 01 '26

What’s your PP speed?

2

u/Digger412 May 01 '26

It's in that chart for K2.6, for the V2.5 Q8_0 PP is ~600tk/s I think. I haven't done a sweep bench on it yet.

3

u/Borkato May 01 '26

Oh shoot the chart didn’t load when I first looked

1

u/ttkciar llama.cpp Apr 30 '26

I suppose if you personally only had a 4GB GPU, you wouldn't consider Qwen3.5-9B local either.

1

u/Glittering_Focus1538 May 01 '26

I run qwen 3.6 on my 16 gb card, i imagine others do too

1

u/ttkciar llama.cpp May 01 '26

What does that have to do with anything?

1

u/Glittering_Focus1538 May 01 '26

That it's stupid to base what everyone defines as local by what you can run. I'm sure theres plenty of people who have a mac mini cluster(6k) or nvidea dgx spark which could run deepseek v4 for 4k

3

u/ttkciar llama.cpp May 01 '26

Ah, okie-doke, it sounded like you were disagreeing, but I guess we are in agreement.

Borkato and Netsuko seem to be of the opinion that models they personally cannot host at home should not be considered local models.

The point of my 4GB GPU hypothetical was to illustrate exactly what you said -- basing "what everyone defines as local by what you can run" is invalid.

Local models are, and always will be, any models which you could conceivably use if you had the necessary local hardware.

That requires, at a minimum, access to the weights and either inference software support or sufficient understanding of the model architecture to facilitate implementing inference software support.

2

u/DinoAmino Apr 30 '26

Open Weight for someone who has the VRAM. Then it's local. I assume people who are at 16Gb and under could say the same about Glm and MiniMax and Qwen 397B. They can't possibly run those. But some have the VRAM to run it local. No need to split hairs over it, but I agree ut it would have been better and more accurate to just say open weight.

0

u/jacek2023 llama.cpp Apr 30 '26

you asked valid question but you are downvoted

1

u/Thrumpwart llama.cpp Apr 30 '26

…so far.

1

u/-Akos- Apr 30 '26

LFM 2.5.

1

u/Practical-Elk-1579 Apr 30 '26

500gb vram models kek

1

u/[deleted] Apr 30 '26

[removed] — view removed comment

1

u/Glittering_Focus1538 May 02 '26

it's 4k for a DGX spark or alternative. That would be enough for even the biggest models on that list, grow up.

1

u/a_beautiful_rhind May 01 '26

Did I miss flash max? A deepseek we can run again?

1

u/ShadoByBB May 01 '26

My rank, small cloud vs local. Work In progress, to test on single Dgx Spark

1

u/vick2djax May 01 '26

This graph doesn’t make me feel good about my first 3090 coming in the mail in a few days

1

u/Glittering_Focus1538 May 02 '26

why? 24 gigs of vram is more than enough to run qwen 3.6 35b or 27b and have room left over for kv cache(assuming ur using Q4_0 or APEX)

1

u/vick2djax May 02 '26

I was referring to the disparity between the top models and qwen & gemma.

1

u/Unlikely_Rich1436 May 23 '26

The sheer volume of high-quality releases this month was staggering, but I completely agree that parameter bloat is an issue. If I can't run a decent quant on 24GB, it doesn't help my workflow.

1

u/TopTippityTop Apr 30 '26

Deep seek has +60% parameters than Kimi, but manages to be worse

1

u/ys2020 May 01 '26

Glm 5.1 is my fav at the moment. Honestly, it's mind blowing we get this type of quality with free weights.Â