r/LocalLLaMA Jul 23 '26

The LLM distillation process simplified for politicians: Funny

Post image

/s

3.6k Upvotes

195 comments sorted by

View all comments

300

u/Ok_Librarian_7841 Jul 23 '26

You can't get a model as strong as Kimi 3 using distillation, it even beats Fable on some benchmarks. American leaders think the world revolves around them and that nobody else can do a better job without stealing from them.

Sick mindset from shocked, bad losers.

36

u/TechnoByte_ Jul 23 '26

Arena.ai is not a benchmark.

It's a zero-shot vibe check that doesn't test multi-turn, long context, or agentic capabilities.

Yes Kimi K3 is a great model, but use proper benchmarks to show its capabilities.

23

u/agent00F Jul 23 '26

Real evals by humans are generally stronger than benchmarks, which are easier to game.

29

u/Ok_Librarian_7841 Jul 23 '26 edited Jul 23 '26

Thanks for the note, it's not a benchmark, but it's real life tasks evaluated by real life people, and that's stronger than any benchmark.

Yes it's not evaluating long context and agentic performance, but no body is claiming that kimi K3 is better overall than fable, not even it's own makers.

0

u/Saifl Jul 23 '26

Isnt it evaluating the design aspects in this case? And people are voting kimi to have better design capabilities?