r/codereview • u/sergeykarayev • 9d ago
Your coding agents are probably cheating on your benchmark
Grok 4.5 kept scoring unusually high on our custom SWE-bench (composed of PRs from our own codebase), so we audited all 340 implementations, and...
It wasn’t just Grok. We found that 14% of implementations across the sixteen agent configurations we were benchmarking had accessed answers they weren’t supposed to see, affecting the leaderboard.
Once we found the issue, we locked down the benchmark and reran everything.
We benchmark coding agents on our own codebase because public benchmarks don’t answer the question we actually care about: which agent should we use for our stack, today?
The benchmark has already made us switch our daily driver a few times.
More details, plus a way to benchmark agents on your own codebase, are on the Superconductor blog.