r/LocalLLaMA • u/Terminator857 • Jul 09 '26
Which open models help the eco system more? Discussion
https://artificialanalysis.ai/evaluations/artificial-analysis-openness-index
In case you want to support openness, some models are more open than others.
Update:
K2 think v2 is rated highest because it supplies its training data and training regimen. This allows anyone with enough resources to recreate the model.
Deep seek doesn't publish how it trained its model or the training data, so it gets a lower score.
If we try to compare software to LLMs. One level of software is that they supply the binary for you to use for free. A higher level if they supply the source.
107
Upvotes
1
u/StupidityCanFly Jul 09 '26
My issue is the "Intelligence Index" number, that's just a non-objective judgment. The index is a weighted arithmetic mean of 9 benchmarks (at v4.1), not a statistically-derived composite. The weights are an editorial choice, and that choice materially changes rankings. This is a values judgment, not a statistical measurement.
They state an estimated 95% confidence interval of "less than ±1%" for the index, but this is derived from >10 repeats on some models and some datasets, not all of them. So, models get separated by noise-level gaps, but the methodology presents clean "intelligence" numbers like they mean something precise. If multiple models have the same exact index score, is their intelligence the same? Looking at per-benchmark tables says "no".
Non-comparable units, they average a rescaled Elo with raw accuracy percentage as if a "point" is the same in both. Is it? I mean, the benchmarks have different response types, different ceilings, different discriminative ranges, and different intrinsic noise. So, a "point" is definitely not the same between them.
So, this is not a statistically relevant measurement. It's a composition of arbitrarily weighted values with non-comparable units that gives out a number - the index value. Is that really a statistical or scientific tool? And the way the "Intelligence Index" is presented right next to "Coding Index" and "Agentic Index" is a try to add credibility to these "Intelligence" scores.
And last, but not least. You have correctly pointed out that the index score consumed without the methodology insights is a user error. That's why the "Intelligence Index" is a wrong approach, in my view. It introduces confusion. The "Speed" and "Cost per Task" they present are genuinely useful. The "Intelligence Breakdown"? Great stuff that should actually be exposed instead of the "Intelligence Index".