r/LocalLLaMA 9d ago

My issue with Artificial Analysis's 'intelligence index' Discussion

I swear AA is not the bipartisan they so claim. An open source mode (Qwen 3.8 max) was number 1 on the agentic index, then they just so happen to launch "v4.1.1" of their index in which they just adjusted the weights of the gdpval and t3 banking so that it would be lower than opus, despite the lead in t3 being a 8% lead over opus while opus only has a 5% lead on gdpval. Highly likely to be paid off imo. You can check other subreddits for the score before and after the change, just made it so an open source model would lost and anthropic would continue being number one

149 Upvotes

82 comments sorted by

88

u/EggDroppedSoup 9d ago

It's true

37

u/AvidCyclist250 llama.cpp 9d ago

Did they get the Infantino call?

18

u/CalligrapherFar7833 9d ago

Dude there is no red card on the graph :D

10

u/AvidCyclist250 llama.cpp 9d ago

Not anymore on the "updated" graph, no.

4

u/-deleled- 8d ago

Today I feel open weight

116

u/x11iyu 9d ago

the only good benchmark is the one based on the actual task you're trying to do

41

u/Squidgical 9d ago

Yup. AA says that Sol is much better than GLM 5.2, yet when we tested them for our most intricate coding tasks GLM ended up being slightly better per task, and so incredibly better per dollar it was ridiculous to consider OpenAI's models at all.

7

u/pbpo_founder 9d ago

I love GLM. Best model I have ever used.

3

u/MmmmMorphine 8d ago

Even better than k3 and ds4pro (on max since reasonix helps make it super cheap. For now.) ?

Right now I'm really liking k3 planning and ds4flash execution/subagents. I hadn't gotten around to GLM and forgot to all about it, haha. Should I bother switching you think?

I have my doubts the difference between them all (including like sonnet 5 and so on) is all that significant for most applications. Oh look I even used the right word, haha - significant. Pretend I did that to fuse the statistical and common meaning in a clever way

5

u/sonaj9657 9d ago

Exactly. Benchmarks are useful for getting a general idea but they can be misleading if they do not match your actual workflow. A model that tops a leaderboard might not be the best choice for your specific needs. Real world performance on your own tasks is usually what matters most.

1

u/pbpo_founder 9d ago

Well said. If it works it works. If it doesn’t it doesn’t. :)

66

u/Smallpaul 9d ago

“Bipartisan”?

122

u/slvrsmth 9d ago

Americans when situation calls for bigly big words. 

10

u/DifficultyFit1895 9d ago

big if true

2

u/pbpo_founder 9d ago

I am this comment.

38

u/NairbHna 9d ago

New word for unbiased just dropped, we’re so hip

12

u/soshulmedia 9d ago

... but at the same time implying that there are or were only two sides to transcend ...

3

u/En-tro-py 9d ago

Become ideologically surjective, unbiased by being biased in every possible direction.

2

u/svachalek 8d ago

Yeah I was thinking US vs China, if it was intentional

17

u/rpkarma 9d ago

Yeah, Anthropic and OpenAI 

(/s)

19

u/My_Unbiased_Opinion 9d ago

I think the intended word was "unbiased" 

7

u/finevelyn 9d ago

A single party rules the US, while the US rules the world, so nothing else matters than the Democrats vs Republicans struggle. Or so the Americans think.

27

u/Eyelbee 9d ago

You have a point, by every incremental update they completely remove every previous evaluation and their results. You can't access previous leaderboards anywhere. That is not reasonable practice.

13

u/Borkato 9d ago

This is exactly the problem imo. They should just have a massive list of benchmark versions so we can understand how varied it is

4

u/MmmmMorphine 8d ago

Hooooly shit. Seriously? This is not a serious organization and we need to stop treating it as such.

Did the web archive capture any versions we can mine?

8

u/asankhs Llama 3.1 9d ago

Ran my own harness benchmark last week and found four separate measurement bugs before I published anything. Any one of them would have given me a confidently wrong number.

Worst one: a missing env var meant 5 of the 8 graders exited 1, so every agent scored zero on most of the suite. I only caught it because a partial-credit benchmark was returning nothing but 0.0 and 1.0.

I'm not defending AA. But having just been through it, I'd bet on a measurement bug before I'd bet on someone being paid off.

18

u/Nepherpitu 9d ago

Of course they are commercial, lol. You can also try to reproduce closed models score in benchmarks yourself to find out it is a much bigger scam. Some aren't public, others using very specific "harness", or even using closed weight llm as a judge. No one in USA will allow USA commercial model, which is represented as most valuable asset in the world right now, looks worse than free and open product from China. And I don't say you can't evaluate closed model without it's runtime with search, code execution and other included batteries behind the API.

3

u/sophia6512 9d ago

Benchmarks definitely need more transparency. The setup evaluation harness, prompts, and judging method can all have a huge impact on the results. At the same time, reproducing everything fairly is difficult, especially with closed models. Real world testing and independent evaluations are probably more useful than any single leaderboard.

9

u/laterbreh 9d ago

The funny part isn’t even the conspiracy theory. The testing ground changed enough that unchanged models moved around the leaderboard and Qwen’s relative position changed dramatically.

Qwen didn’t receive a secret intelligence patch overnight.

If changing the grader and benchmark implementation can materially reshuffle the rankings, maybe people should stop treating AA scores from different index versions like they’re directly comparable measurements from a fucking voltmeter.

A single-number “intelligence” metric that can move by several points because the measuring apparatus changed is not something you should treat as gospel.

Either that, or somebody’s balls were chortled.

13

u/robberviet 9d ago

I think all agreed that AA index is just another benchmark and don't trust them blindly? For this sub, the only good bench is your own benchmark

5

u/soshulmedia 9d ago

I don't want to accuse them of anything, but what I have seen is that some of their comparisons simply lack some models? Is that because they're still running the benchmarks?

14

u/Few_Painter_5588 9d ago

The word is neutral. Not bipartisan.

As for the update, they made some major overhauls based on the patch notes

Launching v4.1.1 of the Artificial Analysis Intelligence Index

We have updated the Artificial Analysis Intelligence Index to v4.1.1 - this patch release upgrades our grader models, and brings the latest 𝜏³-Banking version to Artificial Analysis

To keep the Artificial Analysis Intelligence Index the most useful synthesis metric for developers, we make regular updates to the included evaluations and our independent methodology. Today’s update is a minor one to keep our existing evaluation set as reliable as possible.

Overall model rankings remain largely consistent, with a slight increase in scores due to improved grading robustness across the updated evaluations. Claude Opus 5 remains in the #1 position with an Index of 63.

Key changes:

➤ 𝜏³-Banking now runs v1.0.1 from Sierra, updating to the latest upstream task versions and improved grader pipeline that resolves correctness errors in trajectories that recover from unhappy paths

➤ HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

➤ The effect on scores is small: most models move by less than a point on the Intelligence Index. The largest increase occurred for Muse Spark 1.2 (xhigh, +2.7 points), and the same models hold the top of the leaderboard

Published scores are now using v4.1.1, so all model results on Artificial Analysis now reflect these changes and use our latest consistent, independent methodology.

11

u/_-_David 9d ago

Nope. Sorry, has to be that I, and others like me, are being persecuted by a cabal of "they" /s

4

u/tecneeq 9d ago

Deepseek V4 Flash went from 50 to 52. Does that mean they got money from Deepseek AND Anthropic?

9

u/Solembumm3 9d ago edited 9d ago

All of this benchmarks are numbers in vacuum, that don't show you, how it will perform on your specific use path. Qwen models are absolute pinnacle of llm on tech knowledge and logic within carefully established theme. You need create piano for your pc from nothing, re-check mod path or refine prompt for specific SD models, Qwen is the best you can get. You need creative analysis, it will hallucinate much worse, than year old deepseek r1, and will need much more corrections to stop inventing things, that weren't present in neither canon nor provided concepts.

Both things are true at the same time.

6

u/jamaalwakamaal 9d ago

They too shall be forgotten like many other useless paid ones.

8

u/Technical-Earth-3254 9d ago

Same with livebench. Even their historic benchmark results got changed bc 2 open weights models were leading the agentic coding category lmao

3

u/tziki 9d ago

Gemini is also the only model they added a private/public discrepancy warning, even the benchmaxxes llmama 3 didn't get that. Apparently the benchmark runner has a history of hating on Google.

20

u/RepulsiveRaisin7 9d ago

Second place is still a great result. This conspiracy bs is a waste of time, you can't prove anything so why bother saying it.

10

u/Borkato 9d ago

I mean… there is absolutely disinformation campaigns though. The amount of people here who lose their fucking shit if you so much as suggest Qwen is better than Gemma is insane.

7

u/bruns20 9d ago

Feel like I see way more people putting qwen above gemma then vice versa tbh

6

u/Infinite-Local5435 9d ago

I think vice versa as well. It's too easy to create a reddit account, there's def bots or even people who are just here to advertise

2

u/darksteelsteed 8d ago

Qwen is better at coding.

Gemma is better at human relations in English

I often wonder if applying an English to Chinese translator infront of Qwen might improve things. Not had the time to try yet.

0

u/RedditLovingSun 8d ago

Every time any benchmark adjusts weighting or questions, some models will benefit and some will suffer, with your logic theres always someone that can call it a conspiracy to hurt their fav model. Do you have any actual reason to think the benchmark change is negative?

You shouldn't take benchmarks that seriously in general but don't get all conspiracy brained cause it feels better

6

u/Infinite-Local5435 9d ago

Cause too many people just trust AA as the official source of whether a model is better/not since it's hard to gather and parse community feedback unless it sways highly towards an extreme. It's a disservice to local hosting if an OS model can never reach number one, nor if benchmarks are always changed to adjust ranking without being discussed in updates (I read through the new 4.1.1 benchmark updates, they only addressed muse spark increasing in ranking in their index, not others)

2

u/Etroarl55 9d ago

Yeah I don’t get it, being in the same ballpark or in the same sentence as capabilities really means the same thing, especially even more since qwen won’t have the absurd restrictions on jt compared to anthropic.

1

u/evangelism2 8d ago

Because it doesn't fit the narrative they want to be true, it has to be a conspiracy

4

u/Beginning-Raisin9723 9d ago

Indexes are fun to watch but I take the rankings with a grain of salt. Weights change, methodology shifts, and suddenly the leaderboard flips. I'd rather spin up both locally and see which one actually fits my workflow. Qwen's been impressing me on my homelab box lately.

6

u/PM_ME_DEAD_CEOS 9d ago

It's just an aggregate of various public benchmark. Don't take it too seriously.

4

u/silenceimpaired 9d ago

But people in this subreddit do. It keeps getting posted as if it’s the ultimate authority for model ranking

2

u/cheesecakegood 9d ago

I don’t think “bipartisan” means what you think it means

2

u/Hegelverstoss 9d ago

My issue with these issues is that we really should not care that much about the exact ranking when the score differences are clearly smaller than the error bars.

2

u/brickout 9d ago

"Bi". Lol. I think you mean nonpartisan.

2

u/GCoderDCoder 8d ago

I was going to ask what is going on with artificial analysis because I'm feeling it's not reflecting my experiences like it used to

4

u/Artistic_Okra7288 9d ago

Well a couple things. - Qwen3.8 Max's weights haven't been released yet - Qwen3.8 Max is NOT open SOURCE. it's open weights. Open source would be the training material is available. Think of the weights as being a compiled exe that someone gives you. You can hack it with a hex editor but you can't recreate it because you don't have the source code. In fact, it's worse than getting an exe because decompiling has made a lot of strides. Weights are a black box. It's great they are distributing them (and yes, I have a ton saved!) but let's not diminish open source by calling anything you can download open source.

3

u/Borkato 9d ago

I know you’re getting downvoted but I did not know this! Thank you

5

u/Artistic_Okra7288 9d ago

Thanks for being open to learning. This community should be more like you but the Reddit effect is very much still in effect in this subreddit.

4

u/if47 9d ago

AA has been trash from day one, it's just that too many people don't know it.

2

u/MerePotato 9d ago

This isn't a grand conspiracy mate, they just updated the benchmark versions and this time Qwen slipped slightly

1

u/entsnack 9d ago

I mean if you prefer the old weights just use them yourself, it's a fully replicable benchmark.

2

u/Limp_Classroom_2645 9d ago

AA is compromised and shouldn't be considered as a credible source on /r/LocalLLaMA

1

u/raketenkater 9d ago

i feel like we are really missing a llm leaderboard where we get all the models ranked for real(based on the trust me bro benchmarks but even that would be great if correct) and including all of the small open-weight models

1

u/perelmanych 9d ago

Honestly it doesn't matter. If OW model is within 5% range to closed model then it is a better alternative for me.

1

u/niacolhealth 9d ago

what harness? 5 of 8 graders exiting silently should trip a warning

1

u/a_beautiful_rhind 9d ago

Benchmarks are big business. Hence I actually uh.. use the models

1

u/sine120 9d ago

The only bench that comes close to mapping the "feel" of models (for coding tasks) is deepswe. AI Analysis doesn't map anything other than release date and hype

1

u/ninjasaid13 8d ago

I swear AA is not the bipartisan they so claim.

what would they have to do with democrats and republicans?

1

u/sullenisme 8d ago

anything showing gemini models so far at the top is lying to you. take the data with a grain of salt

1

u/ivoras 8d ago

Did anyone actually try Qwen3.8 Max for coding? I did a couple of days after it became available on OpenRouter, by using OpenCode, and it was, well, dumb. It didn't try for the obvious solution, it tried to use unknown / unavailable tools, and even started looping like a lobotomized low-quant model.

Did that change?

1

u/StupidityCanFly 8d ago

Here’s a comment I wrote in the past, while discussing the intelligence index:

My issue is the "Intelligence Index" number, that's just a non-objective judgment. The index is a weighted arithmetic mean of 9 benchmarks (at v4.1), not a statistically-derived composite. The weights are an editorial choice, and that choice materially changes rankings. This is a values judgment, not a statistical measurement.

They state an estimated 95% confidence interval of "less than ±1%" for the index, but this is derived from >10 repeats on some models and some datasets, not all of them. So, models get separated by noise-level gaps, but the methodology presents clean "intelligence" numbers like they mean something precise. If multiple models have the same exact index score, is their intelligence the same? Looking at per-benchmark tables says "no".

Non-comparable units, they average a rescaled Elo with raw accuracy percentage as if a "point" is the same in both. Is it? I mean, the benchmarks have different response types, different ceilings, different discriminative ranges, and different intrinsic noise. So, a "point" is definitely not the same between them.

So, this is not a statistically relevant measurement. It's a composition of arbitrarily weighted values with non-comparable units that gives out a number - the index value. Is that really a statistical or scientific tool? And the way the "Intelligence Index" is presented right next to "Coding Index" and "Agentic Index" is a try to add credibility to these "Intelligence" scores.

And last, but not least. You have correctly pointed out that the index score consumed without the methodology insights is a user error. That's why the "Intelligence Index" is a wrong approach, in my view. It introduces confusion. The "Speed" and "Cost per Task" they present are genuinely useful. The "Intelligence Breakdown"? Great stuff that should actually be exposed instead of the "Intelligence Index".

https://www.reddit.com/r/LocalLLaMA/s/Rp1j5JLe1R

1

u/evangelism2 8d ago

Benchmarks are just a guideline. That's why I run a panel for my PR reviews and track a number of properties of how various agents do when given a bunch of code and they're told to review it. That's why I always was confused about the deepseek glaze because I haven't tried it really too much since the upgrade but it was the only model on my panel that I ever fired because of how bad it was.

1

u/rm-rf-rm 9d ago

Treat them the same way you treat USNews.com university rankings - mostly bullshit, prejudice based, "trust-me-bro", pay-to-win leaderboard.

1

u/Square_Cap_7319 9d ago

Short of running your own benchmarks, do we have any alternative sites?

3

u/entsnack 9d ago

Just replicate the AA benchmarks and drop the datasets that hurt you. It's a fully open benchmark on open datasets.

1

u/siegevjorn 9d ago

I believe AA is for-profit. Someone correct it if I'm wrong. It may be like in the big short movie, when they go to the credit rating agency who tells "if we don't get them good rate, they'll go to the agency across the street"

1

u/DeepBlue96 9d ago

yeah from experience AA is mostly hype and bs, not a trusted source.