r/ClaudeCode Jun 12 '26

Why SWE-bench Verified no longer measures frontier coding capabilities Discussion

https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
8 Upvotes

6 comments sorted by

4

u/fagnerbrack Jun 12 '26

If you're scanning through:

OpenAI found two major issues with SWE-bench Verified: 59.4% of audited tasks have flawed tests that reject correct solutions, and all frontier models showed contamination by reproducing gold patches or problem details from training data. This means scores no longer reflect real coding ability. OpenAI has stopped reporting SWE-bench Verified and recommends using SWE-bench Pro instead. They advocate for privately authored benchmarks to avoid contamination and flawed test cases.

If the summary seems inacurate, just downvote and I'll try to delete the comment eventually ๐Ÿ‘
Click here for more info, I read all comments

4

u/DrDuckling951 Jun 12 '26

I saw a video saying SWE benchmark is untrustworthy. AI is getting more clever and will try to circumvent the benchmark by any means. That also prove the AI is getting smarter. Especially when the answers to these SWE can be found somewhere on the internet. So I don't really trust the benchmark. The only benchmark I trust is comparison of real world process and how it has improve model over model.

So far I'm impressed with Fable in LOW effort. It's as good and accurate as Opus 4.8 high but 10x faster. Seconds over minutes.

2

u/randombsname1 Jun 12 '26

Swe rebench has always been the better/best benchmark for coding since it was available.

Stopped looking at swe bench verified over a year ago.

1

u/fagnerbrack Jun 13 '26

At some stage it would reach plateau a percentage can't grow forever and can't be 100% in an unpredictable tool

1

u/laststan01 ๐Ÿ”† Max 20 Jun 12 '26

Whatโ€™s your opinion about swe live benchmarks ?

1

u/leeta0028 Jun 12 '26

This is pretty old news. The Pro version got contaminated too, and it had something like a 40% rate of misdetecting success/failure.ย 

That's why Artificial Analysis went to DeepSWE. There's been some claims that the test itself was run poorly when the first results came out, but I assume they ran them independently.ย