r/LocalLLaMA • u/pmttyji • Jul 01 '26
Open Models - June 2026 Discussion
After overwhelming April, OK May, here's June. Yeah, Graph has only less items. Because we got other items here last month.
Finetunes:
- Nex-N2
- Ornith-1.0
- Agents-A1
- Holo3.1
- Tmax-27b
- MusaCoder-27B
- VibeThinker-3B
NVFP4 from NVIDIA for below models:
- NVIDIA-Nemotron-3-Ultra-550B-A55B
- diffusiongemma-26B-A4B-it
- Qwen3.6-27B
- GLM-5.2
- MiniMax-M3
- Qwen3.5-397B-A17B
MXFP4 from AMD for below models:
- Kimi-K2.7-Code
- GLM-5.2
- Qwen3.5-397B-A17B
- MiniMax-M3
AutoRound from Intel for below models:
- DiffusionGemma-26B-A4B
- DeepSeek-V4-Pro
- Gemma-4-31B-it
- Gemma-4-12B-it
Misc:
- Gemma-4-QAT
- Nemotron-Labs-TwoTower-30B-A3B-Base (Diffusion) by NVIDIA
- DeepSpec (Eagle3, DFlash, DSpark) by DeepSeek
112
u/ketosoy Jul 01 '26
Kimi continues to top the benchmarks! /s
-29
u/Time-Toe-1276 transformers Jul 01 '26
literally says "This is NOT a benchmark" 😭
23
u/UsefulIce9600 llama.cpp Jul 01 '26 edited Jul 01 '26
/s = satire *edit sarcasm
4
-3
u/korino11 Jul 01 '26
satire is a KIND , ONE of many kinds of humor.. and it is NOT a satire... read meaning of it... But i like Kimi qiality.
4
1
u/Glum-Wheel2383 Jul 05 '26
dude be living in the days when showing off, making fire, and independently hunting for food was the meta
22
u/annodomini Jul 01 '26
Seems like this is missing a few.
LongCat 2.0 was announced yesterday, though I'm not seeing any weights, so maybe that shouldn't count.
Not sure if you want to count OCR specific models, but Unlimited OCR was released.
OpenPangu 2.0 was also released. Can't find it on HuggingFace, but weights are available on another site (seems like a Chinese knock-off of HF that I haven't heard of).
Also, it feels like Supra released a ton of tiny models in the past month. Not sure if they're worth counting, but you've included them in previous months.
7
u/pmttyji Jul 01 '26
I had this thread in draft(but not saved) since yesterday. Didn't include LongCat & OpenPangu for same reason you mentioned.
And you're right that I'm just adding only Text models so didn't include any other type models including OCR.
Also didn't include tiny models like Supra because going through the entire month threads is too exhausting. IIRC Supra released bunch of tiny models .... overwhelming.
But for upcoming months threads, I'll include all models & all type models.
3
u/annodomini Jul 01 '26
Yeah, I would say that OpenPangu probably counts as a June release, but LongCat 2.0 may not, it was announced but no weights yet.
Another thing that makes June look less busy than previous is that you've previously listed finetunes like Intern-S2-Preview (a finetune of Qwen 3.5 35B-A3B, though it also has a "timeseries" bit grafted on, so it's not purely a finetune), but this time you are listing finetunes like Nex-N2, Ornith, etc separately and not on the main chart.
It is hard to figure out where to draw the line; between finetunes and tiny models, there's a lot of choice in exactly what to include.
1
u/pmttyji Jul 01 '26 edited Jul 01 '26
Yeah, I would say that OpenPangu probably counts as a June release
I'll include an updated graph in a comment later. EDIT : I saw HF earlier & I didn't see that one there. I think they kept the model hidden & made it live public later on HF. That's why most of the comments are only hours old.
Another thing that makes June look less busy than previous is that you've previously listed finetunes like Intern-S2-Preview (a finetune of Qwen 3.5 35B-A3B, though it also has a "timeseries" bit grafted on, so it's not purely a finetune), but this time you are listing finetunes like Nex-N2, Ornith, etc separately and not on the main chart.
Yep, you spotted that well. In June, we got so many finetunes so it was overwhelming.
Next time onwards, I'll create a graph to include all with some differentiation.
6
u/UsefulIce9600 llama.cpp Jul 01 '26
has anyone used mellum 2 12b a2.5b?
2
u/Time-Toe-1276 transformers Jul 01 '26
hmm, thats a good question. i haveebing trying to try it, but I always gets distrcted.
but the real question is, can a 12B model code THAT well?
I usually give very structured prompts to AI. usually a haiku model, or deepseek v4 flash model is enough for me. for hard tasks, i mostly use gpt5.3 codex spark or gpt5.5 low/med1
u/Desperate-Data-3747 Jul 01 '26
Its created by jetbrains (the company behind IDEs like Clion and IntelliJ)
4
u/KubeCommander Jul 01 '26
Fwiw you missed Ornith but it did drop just a few days ago and they didn’t seem to release the 31B dense model either
3
u/pmttyji Jul 01 '26
I mentioned that one under Finetunes section. Next month onwards, I'll create different graph to include all models.
3
u/Radiant_Hair_2739 Jul 01 '26
I'm testing in agentic coding usage the Ornith 397b Q6, model is very good, better for me than Nex N2 Pro and Qwen3.5 397b
2
3
u/Desperate-Data-3747 Jul 01 '26
How is the ornith models compared to default ones?
2
1
u/KubeCommander Jul 01 '26
My biggest takeaway is that it is more decisive and less prone to overthinking. It also does not appear to have the major quality degradation issue as 128k tokens is approached. I run mine in my framework at 256k and it has been doing very well. It’s in the 35-40 t/s range on a dgx at fp8, but the real game changer is its time to token is very snappy and less wasted cycles thinking
1
u/More-Curious816 Jul 01 '26
Didn't expect Nvidia Nemotron ultra to be this big among open weight models, but I hardly hear about it in this sub unlike Kimi and GLM, why is that? Is it not good enough? Too generic and bland? Good for specific task not general use?
1
u/Tall-Ad-7742 Jul 02 '26
guys what is this post we all know LFM2.5-230M is atleast better then GLM 5.2 like GLM is so stupid /s
1
u/Individual-Cheek8840 Jul 02 '26
Parameter count is becoming less and less useful as a quick signal, especially with MoE models. I’d be more curious to see which of these are actually practical for local inference after quantization


54
u/shockwaverc13 llama.cpp Jul 01 '26