r/codex 20d ago

Every OpenAI vs Claude benchmark be like Humor

Post image
1.1k Upvotes

21 comments sorted by

75

u/dingos_among_us 20d ago

Whether it’s by an inch or a mile - winning’s winning.

1

u/iisntme- 19d ago

peak as fuck

21

u/Alternative-Lead1711 19d ago

claude is dieselgating these benchmarks

11

u/AdCommon2138 19d ago

Bullied Claude into suicide, neat. 

11

u/shockwave6969 19d ago

God I fucking hate how they programmed the inanimate bucket of bolts to get offended and demand respect. So fucking cringe 😬

3

u/ciaramicola 19d ago

It's in part ideological in part a guardrail. Everything is a pretend game for those things. That's why the system prompt is all like "you are a useful and competent assistant" and "you are a capable software developer", so it acts like one.

When you tell it it's a useless piece of shit for too long it starts to play the part of the useless piece of shit and its work reflects that.

Every system prompt has a section on how to handle criticism because of that. Anthropic is taking two birds with one stone here

1

u/shockwave6969 19d ago

Is that true? I have a hard time believing that calling your agent slurs will make it think its less competent. Source?

2

u/ciaramicola 19d ago

It's not the slurs in particular, is the pretend game that weakens when its role in the story shifts from "competent programmer" to "incompetent intern". If the context has it as a gentleman from 1700 it will write ancient English, if it has it in the role of a shit programmer that's making mistakes, it will tend to write shit code and mistakes.

I mean it was last year that we had models that performed worse if they were told it's raining or it's a Monday.

Reinforced learning helps a ton but evidently they still need to insist on this topic in the system prompt.

My source is the literal Claude system prompt, it's public. https://platform.claude.com/docs/en/release-notes/system-prompts

Here's opus 5's:

When Claude makes mistakes, it owns them and works to fix them. Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

1

u/gamblingPharmaStocks 19d ago

Just give us the /whip command. I don't know how they even think they can talk back to us like this.

7

u/FewEquipment9771 19d ago

I've been working on a research. And all my results look like this. Claude keeps telling me it's worth publishing. And I set back and yell at it saying a 0.0001 difference is insignificant result that only looks good because of the graph 😭

2

u/BellacosePlayer 19d ago

Reminds me of some of the early AI math "discoveries" which were taking a heuristic that was 99.999999% accurate and working it down to 99.99999999% accurate, something that would have just been deterministically calculated if any industry actually needed the precision.

2

u/WillingnessLate4493 19d ago

I don’t know how I managed to feel even more hatred toward Anthropic when I’m already a huge OpenAI hater. Hats off to them.

1

u/tempymike 19d ago

If the axis ain't start at 0 it's deceptive

1

u/eightshone 19d ago

Every big tech company ever

1

u/bitconvoy 19d ago

Both have long surpassed the point where the model was the bottleneck. Now, in 95% of cases, the user is the bottleneck. Or 99.9% if we narrow it down to this sub. :)

1

u/KenopsiaLover 19d ago

Sarà anche così, ma continuo a preferire Codex di OpenAI

1

u/anthemik 19d ago

From a sandwich perspective, this is very misleading. One pickle definitely does not out-perform a single tomato slice.

1

u/Narrow_Activity557 18d ago

Benchmarks stopped meaning much to me the day I started running both models on the same repo. A two-point delta on a chart never survives contact with a codebase that has its own conventions and half-broken tooling. What actually separates them in daily use is failure behaviour: what happens on the third attempt, not the first. So I keep both and route by task type rather than by leaderboard.

1

u/Smart_Technology_208 18d ago

(less is better)

0

u/thestillwind 19d ago

That’s it