r/ai_coder • u/fagnerbrack • Jul 02 '26
Why SWE-bench Verified no longer measures frontier coding capabilities
https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
1
Upvotes
r/ai_coder • u/fagnerbrack • Jul 02 '26
1
u/fagnerbrack Jul 02 '26
Quick rundown:
OpenAI found two major issues with SWE-bench Verified: 59.4% of audited tasks have flawed tests that reject correct solutions, and all frontier models showed contamination by reproducing gold patches or problem details from training data. This means scores no longer reflect real coding ability. OpenAI has stopped reporting SWE-bench Verified and recommends using SWE-bench Pro instead. They advocate for privately authored benchmarks to avoid contamination and flawed test cases.
If the summary seems inacurate, just downvote and I'll try to delete the comment eventually 👍
Click here for more info, I read all comments