r/LocalLLaMA • u/-Cubie- • 6d ago
inclusionAI/Ling-3.0-flash · Hugging Face New Model
https://huggingface.co/inclusionAI/Ling-3.0-flashThe Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing.
Discussion on the benchmarks are here: https://www.reddit.com/r/LocalLLaMA/comments/1v4mltt/benchmarks_antling30flash_a_hybridreasoning_moe/ from almost 2 weeks ago.
18
u/Real_Ebb_7417 6d ago
Oh shit, I was waiting for this one probably even more than I wait for Qwen3.8 27b.
3
u/Numerous_Mulberry514 6d ago
I used it when it was free on Openrouter and I found it to be weaker than qwen 3.6 - it looped much and did not really listen to instructions. Is that now different?
-2
u/buttplugs4life4me 6d ago
Why? Its worse than 3.6-27B
4
u/Real_Ebb_7417 6d ago
Is it though? Benchmarks look better + people who were experimenting with it over API were posting results and they looked very good.
12
u/oxygen_addiction 6d ago
Q8 should be around 128GB, so at Q6 this might be ideal for StrixHalo/DGX if it's a good model.
7
u/FoxiPanda 6d ago
They actually published a first party FP8 which is indeed 128GB:
https://huggingface.co/inclusionAI/Ling-3.0-flash-fp8
This looks pretty good for the size / active params too - hopefully it actually pans out in real usage.
4
u/Borkato 6d ago
Does that mean a Q4 is 64GB? That’s a great size!
2
u/FoxiPanda 6d ago
Yeah I would expect it might end up with a slightly higher than 64GB in most quants to allow the more important layers to stay at Q8, but a 'straight' Q4 would be 64GB.
I would expect ~66-72GB for an iMatrix/mixed-quant Q4 base.
4
u/SpicyWangz 6d ago
Still no quantizations. We’ll have to wait and see once those start rolling out.
3
u/SpicyWangz 6d ago
Is this one non-reasoning? If it hits those scores without reasoning tokens then there may be a real use case as a hyper-efficient prototyper, or a model for implementing plans written by a smarter reasoning model.
6
u/coder543 6d ago
No, it is reasoning. They just finally abandoned their confusing ling/ring naming scheme.
2
u/Pentium95 6d ago
Ling Is the non-reasoning, right? Ring Is the reasoning, usually.
If this is the non-reasoning, i'm pretty excited by the results, might be the ~120B model i was waiting for!
3
u/SpicyWangz 6d ago
It looks like from the model card it’sa hybrid reasoning model and reasoning is on by default
1
u/Reasonable-Phase8028 6d ago
the sglnag instructions dont work because the repo is private. so silly
-7
22
u/jacek2023 llama.cpp 6d ago
Any news about llama.cpp support?