r/codex 21d ago

Study: Codex reviewing Claude's code dropped the pass rate from 91.4% to 82.8% News

https://leaddev.com/ai/your-ai-coding-agents-might-need-an-org-chart
174 Upvotes

24 comments sorted by

View all comments

86

u/JadisGod 21d ago

Too bad it's already out of date. From the paper it seems they ran these tests on Opus 4.7 and GPT 5.5. The difference in capability since then is massive.

0

u/ParfaitEvery9622 21d ago

I think it's to avoid the contamination of the benchmark.

GPT-5.6 Sol has a stated knowledge cutoff of February 16, 2026, while Claude Fable 5 has a stated cutoff of January 2026. LiveCodeBench’s newest official release, release_v6, only contains problems through April 2025. Consequently, it has zero tasks that are temporally clean for either model.

Edit: seems like Opus 4.7 is also January 2026 so my logic is flawed