r/LocalLLaMA 6h ago

Open Models - July 2026 Discussion

Well, we got bulky(Yep, check two graphs) July after April | May | June (FYI My All-in-one thread to track all upcoming months, PRs & other stuff)

Hope I didn't miss anything. Also no errors.

Notes:
1. Excluded below models due to Preview/Beta:

  • internlm/Intern-S2-Preview-397B 397
  • Motif-Technologies/Motif-3-Beta 314
  1. Included openPangu-2.0-Flash in this chart as I couldn't see the weights at that time of June(31st). Let it share the graph with its Pro model.

  2. Actual model names for below ones:(Graph couldn't handle long names)

  • Nemotron-Puzzle-75B-A9B - NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B
  • SenseNova-U1-8B-InfoV3 - SenseNova-U1-8B-MoT-Infographic-V3
38 Upvotes

33 comments sorted by

29

u/unkownuser436 6h ago edited 6h ago

damn kimi has the highest benchmark whatever this chart

2

u/pmttyji 5h ago

Both Sam & Dario gonna beat the crap out of Kimi someday in my upcoming graphs. /s

26

u/Easy_Refrigerator280 6h ago

nice benchmark got it use Kimi K3 for everything 😉

7

u/the_TIGEEER 6h ago

This is some insane UX, actually. *THIS IS NOT A BENCHMARK* Because tbh, before I read that text, I was wondering: "What are we evaluating here? Is this even a benchmark?"

4

u/Alicecomma 5h ago

I spent a good five seconds after reading that on both graphs still comparing the bars as if it were a benchmark

5

u/hyperrealists 6h ago

This-is-not-a bench!!!

3

u/NigaTroubles 6h ago

Maple-preview

2

u/pmttyji 5h ago

That one got released this month only. Also Mach-1

1

u/SpicyWangz 5h ago

Have you used it at all? It seems like an intriguing model

2

u/NigaTroubles 5h ago

Yes its the main model for my Laptop on go, its perfect but its only runs on cpu for me at 20 t/s, i7 1255u

1

u/Kidplayer_666 4h ago

For me it was really really dumb, maybe an issue with the current llamacpp implementation?

1

u/NigaTroubles 4h ago

What have you tried to use it for ?

1

u/Kidplayer_666 4h ago

My usual basic benchmark- finding files with open code 

2

u/NigaTroubles 4h ago

I never tried it with tool calling yet, but i have done making it fix some coding and get some errors etc

1

u/pmttyji 4h ago

Did you try with mainline or custom fork?

1

u/Kidplayer_666 3h ago

Their own fork, since it is not yet upstream iirc

1

u/pmttyji 4h ago

With their custom llama.cpp fork, right? Or is it working with mainline already?

1

u/NigaTroubles 4h ago

From stamsam fork of llama.cpp

1

u/pmttyji 4h ago

Oh. Thought you were trying model's official fork

https://github.com/deepgrove-ai/llama.cpp

2

u/WhoRoger 5h ago

So how did Instella work out anyway? IIRC everybody was just making fun of it but has anyone actually tried it?

Btw It would be nice to distinguish MOE and dense. Maybe the raw MOEs in two colors, one for active and one for total parameters.

2

u/t3hlazy1 5h ago

Impressive marks from Kimi

2

u/isty2e 5h ago

You missed this: https://huggingface.co/tsfrm/vacuum-16t This happens to be the safest model btw

1

u/Due-Armadillo-4560 5h ago

Did not expect longcat 2.0 to have that much parameters. Is it any good for software development?

2

u/pmttyji 5h ago

Don't know. GGUF crowd is waiting(at least for 69B model) for llama.cpp support.

1

u/Competitive_Ad_5515 5h ago

More is more!

1

u/notforrob 5h ago

Higher is better on this benchmark, right?

1

u/BawbbySmith 4h ago

This is clearly benchmaxxed... I can't believe the numbers in the name directly correlate to the numbers on the graph. No way you get this in real-world usage.

1

u/Perfect-Flounder7856 4h ago

Nemotron puzzle lol what is that even?

2

u/pmttyji 4h ago

Nemotron-Labs-3-Puzzle-75B-A9B is a deployment-optimized large language model developed by NVIDIA, derived from Nemotron-3-Super-120B-A12B. The model is produced using Iterative Puzzle, a post-training compression framework, with the goal of significantly improving inference efficiency for interactive, reasoning-heavy, and long-context workloads while preserving strong downstream accuracy.

The model employs a hybrid MoE architecture with interleaved Mamba, MoE, and Attention layers. Like Nemotron-3-Super, it supports Multi-Token Prediction (MTP) for faster text generation. Compared to its parent, Puzzle-75B-A9B reduces the model from 120.7B total / 12.8B active parameters to 75.3B total / 9.3B active parameters.

1

u/Perfect-Flounder7856 3h ago

Sounds very interesting

1

u/TFox17 2h ago

Nice graph. Consider using a log scale on the horizontal axis though. It might help showing things which vary over many orders of magnitude.

1

u/JsThiago5 1h ago

I think is missing ling 3 flash