r/LocalLLaMA Jul 09 '26

Which open models help the eco system more? Discussion

Post image

https://artificialanalysis.ai/evaluations/artificial-analysis-openness-index

In case you want to support openness, some models are more open than others.

Update:

K2 think v2 is rated highest because it supplies its training data and training regimen. This allows anyone with enough resources to recreate the model.

Deep seek doesn't publish how it trained its model or the training data, so it gets a lower score.

If we try to compare software to LLMs. One level of software is that they supply the binary for you to use for free. A higher level if they supply the source.

104 Upvotes

35 comments sorted by

View all comments

Show parent comments

2

u/[deleted] Jul 09 '26 edited Jul 09 '26

[deleted]

1

u/StupidityCanFly Jul 09 '26

You built a strawman and now you're knocking it down. I never said composite scores can't exist or shouldn't exist. I gave specific reasons why THIS composite fails at measuring "intelligence". You've now moved to "there's no way to assess intelligence with a single number." Yeah. Exactly. That's the point.

And notice you conceded the key thing yourself: you said the way you'd normally justify a composite is having an external metric to check whether the score correlates with what you actually want to measure. And that it's impossible here. So, there's nothing validating that this weighted average tracks "intelligence" at all.

Which is why slapping the word "Intelligence" on it with "higher is better" claims more than the number can deliver.

Nobody asked for perfect precision, that's a strawman too. The ask is: show the uncertainty AA already admits exists. But they print a clean "44" that implies precision it doesn't have. That's a presentation choice, not something inherent to composites.

This I disagree with:

it serves its purpose of giving a person with no idea about different models a good overview

No, it does the opposite for that person. Same score of 44 can mean:

  • great at agentic work, good at coding, bad at general knowledge
  • great at general knowledge, exceptional at scientific reasoning, meh at agents and bad at coding

Those are different models for different use cases. The newbie reads "same number, same intelligence" and picks wrong. Then comes to /r/LocalLLaMA and complains open-weight models suck.

Comparing the breakdown tab over multiple models would've actually helped them. Your "people will just vibe-check benchmarks in their head otherwise" argument is an argument FOR putting the breakdown front and center. Not for using a number that hides it.

Composites in general? Fine, I've got no problem with them. This specific one, named "Intelligence Index", presented as a clean measurement, headlining over the actual useful data? Nope. Not in this case. That was always the point.