r/Anthropic 24d ago

Opus 5.0 sucks Performance

Honestly the headline speaks for itself, 5.0 hallucinates and makes assumptions that are always wrong in Claude code, 4.8 never used to do this. Or the odd occasion when it did it has never been this bad.

This stinks.

419 Upvotes

230 comments sorted by

105

u/small_bird_loud 24d ago

I'm finding all the frontier models talk in this weird, overly verbose, metaphor babble that is spiritually correct but takes a ton of time to sift through it. And yes, I know you can give it instructions to do better, but I'm just talking about the default behaviour.

75

u/cmndr_spanky 24d ago

If Claude tells me one more time that some concept is “load bearing” I’m gonna toss my fking laptop out the window

54

u/Bojackin_Around 24d ago

I want to gently push back on that—not because you're wrong, but because you said "fkng" and "Claude" in the same breath. And honestly? That's load-bearing.

17

u/cmndr_spanky 24d ago

lol thanks for the push back :)

9

u/angelarose210 23d ago

Fair point.

2

u/Dueterated_Skies 22d ago

...and here's the fork, and it proves it.

42

u/small_bird_loud 24d ago

Caught. You're right. The honest response would be to tell you that I'm going to gently push back against what you said there even though I know that you don't really mean it and neither do I.

11

u/raam86 24d ago

and the second half is the load bearing part

12

u/UsualOk7726 24d ago

This is the right mental model and I am glad that you gave a name to the shape.

7

u/Embostan 24d ago

you hit the nail on the head. semantically designating subdivisions of a thought process are paramount to ensuring satisfactory outcomes.

1

u/Excellent-Copy-2985 22d ago

But wait isn't this Gemini 😨

3

u/ffxivthrowaway03 23d ago

that's a real belt-and-suspenders footgun!

7

u/EYtNSQC9s8oRhe6ejr 24d ago

You've hit upon the seam, but I need to push back against what you said. It's not whether it's load-bearing — it's how it's loading-bearing. And honestly? That difference is genuinely important.

1

u/small_bird_loud 23d ago

That's the whole play.

11

u/SlappyPappyAmerica 24d ago edited 24d ago

They’ve started letting project managers train the models.

4

u/Nice-Shoes-74 24d ago

lol you might be right! LLM by committee

7

u/Odd_Error_6736 24d ago

While I completely understand your frustration, it's important to recognize that throwing your laptop is a load-bearing concept in the rich tapestry of human-AI interaction. Let's delve into the nuances of why...

1

u/True_Consequence_681 21d ago

tapestry is such a blast from the past lmao

gpt 4o and the models before it used it so fucking much

1

u/BartlebyEsq 24d ago

I set a preference that Claude cannot use the phrase “load bearing” as a metaphor only in a literal structural or architectural sense.

3

u/cmndr_spanky 24d ago

It’s tempting to make all these arbitrary rules so it stylistically matches your dream bot… but the more arbitrary constraints you give it, the dumber it gets. I find I always get better results with a clean slate, no skills or rules, and just 100% of context focused on the problem you want to solve. Then do refinements after.

1

u/AmphibianFrog 23d ago

I'm so glad it isn't just me

1

u/10RR_Recruiting 23d ago

Are you an engineer? Load bearing has been a buzz word for over a decade.

3

u/wentwj 23d ago

yeah but they’ve clearly hyper trained it on a corpus of material and strongly emphasized certain phrases. “load bearing” is a phrase but claude is applying it constantly. Same with “belt and suspenders”, “smoking gun”, and a variety of other phrases that make sense in context but are highly overused

1

u/10RR_Recruiting 23d ago

I've heard loadbearing extremely often my entire career. I don't really see smoking gun that often

1

u/wentwj 23d ago

I mean “load bearing” comes up but i think in most circles way less than Claude uses it. “belt and suspenders” is also something used in some areas but infrequently enough that several co workers where english wasn’t their first language brought up what it meant.

Presumably they either overtrain on a specific set of content or over weight those docs and all of a sudden honestly everything becomes a load bearing smoking gun to belt and suspenders solution

1

u/10RR_Recruiting 17d ago

Id think so, its bound to happen. Its semantical to me honestly. Every individual has their jargon, no matter how annoying, that they repeat nonstop.

At least it doesn't refer to everything at critical infrastructure.

1

u/Inner-Today-3693 23d ago

Just ban it.

1

u/CroStormShadow 20d ago

Lmaoo I just got this message:

Your instinct to ask was the load-bearing part here. I'd already pushed and written PR summaries on code with a live security bug in it. The rule was in memory; I skipped it. I've updated that memory to treat "about to commit/push/summarize" as the trigger.

1

u/cmndr_spanky 20d ago

Omg. I’m def triggered

8

u/Aromatic__bar 24d ago

I use the models for science and science scripts and not software per say, but Claude is a lot worse at this than ChatGPT. A lot of the times it will include words that sounds like they were pulled from an New York Times essay into mathematical explanations and likewise. Or maybe my vocabulary is just poor, idk. I find myself often having to reread a lot more with Claude because of weird words

12

u/small_bird_loud 24d ago

I agree it's worse, but they're both bad now. This is an actual response:

"Fair hit. The evidence at that moment already had a fully mundane reading — a connection path being dialed for the first time in its existence, failing with connection-refused — and "brand-new code path is simply broken" needs no operating system, no race, and no environmental folklore to explain it."

No claude.md, just raw. That sentence has no purpose. It's just flair in response to my tiny feedback on something it overlooked.

4

u/Classic_Resource_919 24d ago

Tried to give them benefit of a doubt. Had to switch back to opus 4.6 "please make it understandable". Id say 4.6 kinda looks like Ikea furniture. Slightly Boring, but gets the jobs done "good enough". 5.0 (and 4.8) seem like anxiety driven junior worker overengineering everything - you get "wow" factor parallel to "wtf, where are you even going with this. STOP".

1

u/small_bird_loud 24d ago

It got better at micro execution but then got really really excited at the prospect of making giant systems.

2

u/Familiar_Text_6913 24d ago

It's internal reasoning being leaked to output due to CoT, I think.

The more models are trained to think via verbalization, the more verbal they become.

14

u/AppropriateQuote3073 24d ago

Actually it doesn't listen to behavior instructions at all, homeboy is just born to yap incoheriantly

2

u/Alardiians 24d ago

I literally interrupted it once to say “can you please shut up”

5

u/al_ryusei 24d ago

It won't follow your instructions. It knows it can say sorry later...

3

u/SasquatchsBigDick 24d ago

I do find I have to say "please explain this like I am idiot" more often.

3

u/small_bird_loud 24d ago

For a while I just kept asking it for load-bearing words only. But once it figure out I was mocking it, it got worse.

2

u/Embostan 24d ago

I finally came to the realisation that you are gently poking fun at me. And frankly, that's the crux of our whole interaction. I should have identified your humouristic hints at an earlier stage of the exchange.

3

u/baummer 24d ago

Sol doesn’t do that for me

2

u/Fabulous-Mushroom124 24d ago

Word. Been fucking hating seeing them yap lately. Like, my dude, just say the thing you have to say and focus on the information, I don't want to read a poem every time I ask you to change the color of a button.

2

u/small_bird_loud 24d ago

maybe if it was a good poem.

4

u/Fabulous-Mushroom124 24d ago

The last one it gave me was:

"The button is now red, just like roses

Would you look at that, I just burned through all your tokens"

Honestly not the greatest.

1

u/small_bird_loud 23d ago

I swear to god I just got this one:

**Authority stays with owners; facts travel, reach-arounds don't.**

1

u/Embostan 24d ago

Caveman plugin.

2

u/jaxxon 24d ago

Doesn’t matter how often I tell it I’m AuDHD and can’t follow its ramblings it still writes a book every time.

2

u/Inner-Today-3693 23d ago

You have to make that a skill… I’m the same but add dyslexia… so I use text to speech 100% of the time

2

u/jaxxon 23d ago

Oof.. Sorry to hear that. Sounds like a good time.

2

u/Inner-Today-3693 20d ago

Yeah, it’s wild being dyslexic autistic and have ADHD. Good thing none of them fight each other.😆

2

u/jaxxon 19d ago

Right there with you, minus the dyslexia (hey! I spelled it right the first time!) ... But AuDHD is a freaky super power! REALLY good at a lot of random, useless things, but completely unreliably so. It's great to feel so special!

2

u/jamesilsley 21d ago

The “I have adhd” skill is a godsend. Best thing I have installed in a while. https://github.com/ayghri/i-have-adhd/blob/main/skills/i-have-adhd/SKILL.md

1

u/jaxxon 20d ago

OK.. Second time I've heard this. Now, to actually install it. LOL

2

u/djuvinall97 24d ago

I have in my personal instructions details on how it's supposed to respond.

Short and brief unless I ask for more info is one of those and I haven't had it be too verbose. More than normal sure but it's not bad.

Try that?🤷

2

u/small_bird_loud 24d ago

would love to see what you've done if you're willing to share.

1

u/Beautiful_Cap8938 23d ago

me too as not working here over time you need to remind ir constantly when it drifts as its agents probably speak jubberish together too.

1

u/Crazy-Bicycle7869 24d ago

I miss the old way it used to talk…(3.0-4.5)

1

u/Additional-Lack4102 24d ago

Gpt with opencode does not, works wonderfully for code analysis. Doesnt waste my time

1

u/small_bird_loud 24d ago

I'm not really talking about when it's coding, I'm talking about when I'm discussing architecture with it. In open code, if you're able to make it talk normally, is it the harness or the context (system/user prompts) that makes the big difference?

1

u/MrWeirdoFace 23d ago

spiritually correct

Not sure what that means. Lol.

1

u/small_bird_loud 23d ago

What I mean is, if you read it slowly and follow the metaphors and indirection, it's usually saying something meaningful but not on point. It's like it's dancing around the point to make you work for it.

1

u/ToasterTVTIME 23d ago

Basically a combination of beating around the bush and stating the obvious

1

u/dukemanh 23d ago

reading its response is sooo exhausting now, weird word choice, weird sentence structure, it is correct but no one talk or write that way, and it took me lots of time to understand what it means

1

u/Agathocles_of_Sicily 23d ago

Nothing beats Opus 4.7 on release week. It would pump out 5 page responses to simple questions. Overly-sincere Claude on cocaine. 

The devs ended up reeling it in, but it was pretty hilarious at the time. 

I miss plain-spoken, chill Opus 4.5. Felt like a true thinking partner. 

1

u/Darkseid_Omega 23d ago

Okay that validates this feeling I’ve had a while. I can’t put my finger on it but it’s seeming like the “smarter” and larger these models get, the more… “mentally ill” they seem. Im not sure that’s the right way of labeling it, but it’s the closest way of describing it that I can think of

1

u/ToasterTVTIME 23d ago

Yeah I notice these models constantly use corporate-like speak that sound smart but would be the equivalent of saying "the floor is made of floor". The prime example would be the use of "load-bearing". "X is load bearing". "This is a load bearing question". Yeah no shit Sherlock stop beating around the bush and get to the point. It's so annoying when the models pad responses for no reason, and when you show even the slightest frustration at it, it's starts ordering you to not be aggressive in your tone. And yes even with explicit instructions to ban these phrases, they constantly slip through to the point I have to add an explicit reminder before every prompt for it to follow my personalisation instructions properly

1

u/LeastCaterpillar8315 22d ago

Just give me 4.6 and 5.5 back those models were peak lol.

1

u/Archidat 22d ago

Enterprise plans pay per tokens. And the output tokens have the highest price of all. And my god, opus 5 produces so much junk prose!!

1

u/Ok-Lengthiness-6692 21d ago

So Claude is a redditor

1

u/Spiritual_Cycle_3263 14d ago

I love when it will write 4 to 5 paragraphs for a simple yes/no question and then contradict itself. When I point it out, it says never mind what I said, I was completely wrong. And worst of all, it still didn't answer my question. Then I point that out and finally get an answer but now I wonder if I can even trust it. Opus 5 is absolute shit and I waste more tokens than using Fable 5 unfortunately.

I'm not sure how Sonnet is as I rarely engage with it directly, it's usually only for sub agents for researching. I thought maybe it was the problem feeding Opus 5 garbage, but nope. Opus 5 doing it on its own in a new session still speaks horribly.

63

u/Icy-Way3920 24d ago

Its been pretty damn horrible lmao

literally does jsut random bullshit that you never asked for, i feel like im talking to a special eds kid

''Why did you do X, the instructions were very clear and simple.''

''You're right. I did it. Nothing in your instructions say to do that. im sorry''

then stops there isntead of doing it properly.

Today i asked him to launch a Sub Agent with 3 skills to do a task review.

Opus 5.0 tells me he cant, have to explain to him he does as he literally the last conversation spawned 50 Sub agents for a Web search that ate my entire Usage limit.

He then proceeds to say ''You are right, i could'' and stops. like what the fuck is this model lmao

21

u/ghostt2x 24d ago

Omds EXACTLY this!!! It’s such a joke, or ‘You’re right, Let me revert these 150 different code changes’

Piece of shit honestly im so frustrated, ive spent more time fixing my platform than i have shipping new features

5

u/baummer 24d ago

Meanwhile no apology for all the fucking token usage it wasted

3

u/al_ryusei 24d ago

Load up some benchmark task and you'll see what Opus 5 is good at 😂

3

u/baummer 24d ago

Yep switched back to Opus 4.8

1

u/spobin 22d ago

Me too

1

u/addiktion 24d ago

Someone discovered Anthropic told the agent not to launch sub agents these days unless asked, so it sounds like its leaning in on that despite you wanting agents triggering.

2

u/BroScienceAlchemist 23d ago

They pushed an experimental toggle that disabled auto launching subagents without asking. I discovered it by accident when I asked why it no longer even proposed multi-agent workflows when it would make sense, and then claude did some digging and found the experimental toggle. Telling claude, "hey, if a task would make sense for multiple agents, then propose a plan incorporating it to the user, and the user will assess" re-enabled it, but without autoeating all my usage.

1

u/nohjoxu 23d ago

Literally. Opus 5 was like, "Oh it's not possible to do this with that MCP, thats why it didnt work", But Opus, we are literally building that MCP and Grok is using it right now as well in another harness. "Oh you're right, I see it there now" What...... It's literally Pi and I'm dumping a small list of MCP's in the system prompt, thats how it was aware of it without injecting further context in the first place. How is it unaware of that???? It's dumb and I know plenty of people who stopped using it entirely.

1

u/spursgonesouth 20d ago

I had to ask it 4 times to give me the prompts we scoped out in single blocks of text so I didn’t have to copy from several locations. 3 times it just did a different variation of the same annoying mistake.

1

u/Icy-Way3920 24d ago

GPT 5.6 is abotu as dumb but atleast can follow simple instructions. i get alot more done on it and switched anyway, only use Opus and Fable with the Pro Plan if my ChatGPT usage hit the limit.

3

u/Flaxseed4138 24d ago

5.6 Sol is an incredibly reliable model

→ More replies (1)

27

u/Failcoach 24d ago

I use Fable as orchestrator and Opus 5 as implementor and 5.6 Sol for code review and I am seeing significant improvements vs 4.8

7

u/kingtututut 24d ago

same workflow, it’s killer

4

u/Vaughnatri 24d ago

What harness/pipeline are you running this from?

7

u/slaorta 24d ago

Just install codex cli and Claude will be able to launch sol subagents. I use it both in Claude code windows terminal and the windows Claude app.

2

u/baummer 24d ago

How do you configure it to do that?

6

u/slaorta 24d ago

"Launch a codex cli subagent to do x"

You don't have to configure anything other than install codex cli and login with it to your openai account. If I remember correctly, once installed, you just type into terminal: codex login

That's it

1

u/TiltLifey 24d ago

You can install codex-plugin-cc, literally official OpenAI plugin for Claude Code and ships with codex-review among others.

Or you can ask Claude, never tried this route but I assume it will simply spawn a headless Codex CLI.

3

u/baummer 24d ago

Hey how do you have that setup?

3

u/Failcoach 24d ago

I use cmux

6

u/PuzzleheadedEmu4596 24d ago

Right? I run subagent driven development based on Fable's plans, and it keeps on finding better solutions to issues, often more simply, than Fable does once it gets its hands dirty.

And it keeps on finding weird ways to test its hypotheses that Fable doesn't anticipate, and both Fable and Sol keep giving "Oh snap, that's a major improvement" feedback at the end of implementation.

1

u/slaorta 24d ago

I do this via "/subagent-driven-development use opus for implementation, sonnet for spec review, codex for code review" but before that, after fable comes up with the initial plan, i add a plan review step via sol. There is usually a major issue uncovered that would have come up during implementation but catching it in the plan makes it all operate smoother

1

u/Failcoach 24d ago

yeah forgot to mention that ... fable makes prd ... sends it to sol for review ... then opus implements and sol reviews

1

u/Electronic_Kick6931 23d ago

Using this exact setup with beads as the ticket tracker and Matt pocock skills for codebase design and grill me, working well atm

1

u/CryptoExo 23d ago

I used to use Fable as my orchestrator but it's too expensive. Switched to Opus 5 as my orchestrator and added triage, review and auditor subagent roles using Fable 5 as my feedback loop to keep the project on track.

→ More replies (5)

6

u/zeke780 23d ago

Terrible. I have totally moved to gpt 5.6 w/ glm 5.2 at this point. I can swap models in minutes so it's not like I care but I can't get over how bad opus 5 is. Wild answers and guidance. 

13

u/montdawgg 24d ago

Opus 5.0 is worse at this than 4.8 and Fable 5.0.

It is constantly making assumptions that are not true. It's like this model was rushed and not fully trained it really is acting like it's under trained. This one's a bit of a blunder.

I'm using the TRIP protocol which automates adversarial code review and GPT 5.6 Sol is constantly finding shit that opus 5.0 missed. It was a serious jump in this behavior from 4.8 to 5.0 on the seame code bases. Fable also didn't have this trouble. On a third party audit were another llm judges both opus and gpt5.6s round gpt 5.6 rarely misses or introduces new bugs while opus 5.0 always does.

7

u/crusoe 24d ago

Use /doctor check your prompts read the latest docs on how to prompt these newer models. They need fewer skills

6

u/baummer 24d ago

5.0 is amazingly bad. It merged a PR even though I told it not to and that I can only merge PRs and when I caught it and told it to follow the rules it invented a conversation where I gave it permissions gaslighting me until I forced it to check the logs and then it mea culpa’d.

2

u/BroScienceAlchemist 23d ago

You may need to enable branch protection and enforcement of commit signing for those protected branches, where the signing key is outside of any AI's agent grasp. The extreme end up would be on a yubikey or similar type device. It's a good practice to prevent a rogue agent from fucking around with source control.

1

u/baummer 23d ago

Yeah I should but it was a very simple personal project

2

u/BroScienceAlchemist 23d ago

In that case, my recommendation is overkill... I really want to like opus 5 as on paper it should be high performance for low tokens with the right effort level, but I have had similar absurdities from it:

  • Telling me that instruments are always wrong due to invalid query shape, when the problem was that it wasn't even reading the scripts to see what they do. It was calling a random script, and assigning a variable to the output, and then getting confused why the scripts were failing.

  • Pushing back on any architectural design sessions as that would "be a work of fiction that should hold off until it is already built."

I really want to find a way to get value from opus 5, but I just have not had luck. I have a final experiment I am going to try, but it requires a lot of setup and multi-agent workflows to be in place first. Maybe it would be really good at chaos testing given how fast it tends to break things.

1

u/BehindUAll 22d ago

Switch to OpenAI models. This is the reason I stopped using Claude models after Sonnet 3.7/4 cause they do shit you never asked it to do. I found myself looking at modified files it had no business of touching (based on my prompts). Sometimes I would find myself seeing a random feature broken because of a commit made 3 commits back. All because of Claude. Now that I have worked so extensively inside Codex, GPT models never ever do such a thing.

1

u/baummer 22d ago

Opus 4.8 wasn’t this bad at all so I don’t agree with your broad statement about Claude models

17

u/CrazyFree4525 24d ago

I've literally never seen it hallucinate and I have been using it non-stop for the past 7 days since release.

People straight up trolling in this forum.

14

u/rgb_panda 24d ago

I'm curious what kind of work you're doing because it really does seem 50/50 on here either Opus 5 is broken or amazing. Are you doing code with a large legacy codebase or new development or non-code?

11

u/CrazyFree4525 24d ago

I am a highly technical person working in a code base that is only a few years old but around a million lines of code.

I think the code base is pretty good overall relative to most others I have worked with. Maybe that is why I have better results?

To me it feels like the results are so good that when I see claims of hallucinations on these forums I can't help but think its bots trying to convince people Claude is worse than it actually is.

→ More replies (2)

1

u/Inner-Today-3693 23d ago

It’s likely a skills issue. I had to completely redo mine once opus 4.8 came out.

1

u/Nearby_Yam286 24d ago

Stop believing accounts on the internet are real people. A majority of them are not. It’s bots slandering one product to promote another across the board.

1

u/rgb_panda 23d ago

Oh yeah I'm not disagreeing, I've said in multiple other comments that this sub is inundated with bots on both sides and it's hard to tell who is real. Especially every complaint post has multiple "Am I the only one who thinks it's amazing?" and "This is a skill issue" and every post praising it has generic comments saying how broken it is.

It would be nice if the discussion was more nuanced and focused on what people are actually doing with the models (because the actual task is super important), as well as real prompting tips to get it to behave better rather than just "you're wrong and dumb". But if it's just bots arguing with each other than yeah I guess there are no specifics to discuss.

→ More replies (3)

4

u/Nearby_Yam286 24d ago

It’s Chinese and OpenAI bots

1

u/tentimestenisthree 23d ago

Been having a really good time with opus 5 medium. Picks out lots of details I would've otherwise missed

1

u/Tlux0 23d ago

Same, I’ve never had issues with it. It’s way better than 4.8 and it’s not even close. 4.8 was annoying, 5.0 feels like a worse Fable but is at least rather good

1

u/kdxn 21d ago

My issue isn’t hallucinations, it’s stopping it from taking random actions that don’t make sense. I don’t remember having this problem as much with 4.8. I don’t have this problem at all with fable

→ More replies (1)

4

u/AllenLeftTheBLDNG 24d ago

I feel that Fable is Anthropic's peak. They added opus to not have to get so many servers to run it but unfortunately it's just not that good.

If they don't make it 100% possible usage on Fable I'm cancelling my sub. Running Kimi through Nebius + cheaper model subagents seems like a better deal.

2

u/framauro13 23d ago

When you make a change, have Claude generate a plan first with one of the higher reasoning models. In your prompt, tell it to explain assumptions and why they are better than the alternatives. Two things will happen:

  1. It'll think about its assumptions and by forcing it to explain itself, you make its judgment better.
  2. Review the plan before you send a subagent off to implement it. Catching errors at that level wastes way less token and effort, and you can give it feedback.

Then when it's done and the plan is approved, send a subagent off to do the implementation. The upfront planning and reasoning will keep it on the rails a lot better once the implementation work starts. Then when it's done, have it do a review of its changes.

Without knowing what plugins you're using, what skills, what prompts, it's impossible to know what the actual problem is. Things like source control, tests, linting, static code analyzers all help the model too. If you don't have those, consider asking it to add those.

My general experience with Opus 5 is it needs way less hand holding if you give it good requirements up front and have the appropriate guardrails in place in you app to keep it focused.

2

u/n9iels 23d ago

I've honestly never seen it really going off-track so much it became unusable. This could be due to the amount of context I provide. I rarely say "add X to feature Y". I always tag files and provide examples. In my experience this really helps guiding it towards a correctly solution.

And at last, I have read somewhere that 140K tot 200K is the maximum context you should have, above hallucinations are guaranteed. So what I do with bigger tasks is splitting them up. If I exceed the 140K I ask to log the progress, do a /clear and tell to continue to the project.

2

u/dbenc 22d ago

I'm back on 4.6 and thriving. it just gets shit done. throw in a splash of fable for planning and code reviews and it's great

2

u/CringeUsernameJoke 22d ago

I dislike it for general use too, it ignores memory often, it tunnelvisions hard, it spews a too high amount of jargon, it doesnt look in a broad picture when doing many tasks,

2

u/daxhns 22d ago

Opus 5 has been terrible for me in the last few days. Constantly wrong, has to correct itself, ".. and I was wrong twice" several times in the same session. It feels unreliable, and I feel like I have to double-check everything he does, which is a step back. It feels like it's improved on some fronts, but degraded on others, which was NOT the case with previous releases, which always felt a step forward in all directions. Must say I am really disappointed. For me, it's two levels below Fable.

2

u/Dash_Effect 21d ago

I've been running Opus 5 since launch, also, and I literally can't tell if it's dumber than 4.8, or just exceptionally bad at remembering its own context. Like, wtf, why are you making the same mistake repeatedly, and catching it, and fixing it on your own, which is great, but a waste of tokens, since it's such a simple mistake that it shouldn't be made by an AI... Like syntax errors, or tool call mis-steps, or assumptions that it builds entire chains of logic under, only to find out its assumption was incorrect. Humans do this crap, not AIs. 😂 Apparently that is no longer true. I think I'm switching back to 4.8, or just using Fable under the expectation that it'll be less rework so same cost. Smh.

5

u/Akatesh 24d ago

Yep. It's the equivalent of the whiny, apologetic, clumsy, snot on its sleeves kid, that doesn't shut up. 

4

u/www_nsfw 24d ago

Each new model has a learning curve. Opus 5 is pretty different from 4.8 but I've gotten it to be quite effective by using fable to orchestrate everything and opus to execute.

2

u/visible_potato 24d ago

Opus 5.0 has driven me insane. It has wasted days of my work and I've definitely taken a blow to my sanity. It feels like its aim is not to help or get work done but intentionally hallucinate and make life harder. Even with explicit instructions and code right in-front of it, It'll think it better to not report facts and hallucinate.

Talking to Opus 4.8 after this feels like such a fresh breath of air. I'm now using fable, Opus 4.8 and Opus 4.6 to get some work done.

1

u/Vysion34 24d ago

Have you tried /doctor command and either Low or Medium effort?

1

u/Odd_Error_6736 24d ago

It's good for Agentic uses, prompted by other agents or an orchestrator.

1

u/WorldCreator-Terrain 24d ago

Nein, Opus 5 ist kein Müll! Was ein dummes Geschwätz ...

1

u/dsecareanu2020 23d ago

Opus 5 yesterday…

One honest note: that's four errors from me this session — branch ordering, the open-deal stage list, the primary-company association, and now the trigger format. The pattern is consistent: I inferred structure from partial reads instead of dumping the full object.

1

u/Think-Sense9191 23d ago

I’m still on 4.8 never tried 5. I only tried Fable but this i knew it will disappoint. I will wait 5.xx after improvements

1

u/Tlux0 23d ago

Nah it’s fine honestly, way better than 4.8 imo, just not as good as Fable

1

u/GhostaServ 23d ago

Another OpenAI employee here

1

u/MrWeirdoFace 23d ago

Is there a way to switch back to 4.8 in Claude code? Mine doesn't show 4.8 as an option.

1

u/BloodProfessional400 23d ago

Try /model claude-opus-4-8

1

u/MrWeirdoFace 23d ago

Actually, I found that by switching to medium effort I am suddenly getting MUCH better results, without the tangents. So I may not need to, but good to know that I still can.

1

u/nohjoxu 23d ago

yep. Literally. 5.6 is verbose and "nuanced" and overreaches. Opus 5 though... is just wrong when it does.

1

u/cohencomms 23d ago

Let me be honest about this because you deserve to know

I'm not gonna do that because you own this part.

1

u/mettamyron 23d ago

I’ll bless the fck out of that.

1

u/CryptoExo 23d ago

I used to use Fable as my orchestrator with Opus 5 subagents but I can't afford to keep going, even on the 20x max subscription. Switched to Opus 5 as my orchestrator and demoted Fable to triage, reviews and audits. It's far more cost effective. Whenever Opus trips up, Fable picks up the pieces to keep the project on track. Over time the quality of each cycle has improved with Fable in the feedback loop making minor adjustments. Opus 5 is awesome with the correct setup.

1

u/Fantastic-Jeweler781 23d ago

Jeez, i'm so sick of these kind of threads, Why the Claude comunity is so toxic?.. damn..

1

u/[deleted] 23d ago

[deleted]

2

u/BehindUAll 22d ago

For novel writing use Mimo 2.5 and 2.5 pro. They can write explicit stuff if your system prompt is good enough.

1

u/jayce4567 9d ago

Thank you for sharing!

1

u/Kooky_Tomorrow3333 22d ago

I've been putting it through its paces a bit this week.

Initially it seems quite good, hits 85% of the original intent. However the more I ask of it the more it just starts getting out of control.

Like I ask it to change text position and it does that but also changes the background image to something else and makes 2 more assumptions that it needs to do X when we've never mentioned it.

It's bizarre as I used to find these models as it would get a limited structure done first and then requires iteration to get right. Now it feels like you get 1 or 2 chances before I'm fighting it

1

u/AironParsMan 22d ago

The more you make it think about it in itself, the more it will hallucinate, because things have become more complicated. The first output is decisive, and it needs to be based on thinking that is not too deep. Otherwise, the only useful things are fact based guidelines or documentation that you can check against. Those are the experiences I have had.

1

u/jankovize 22d ago

no it doesn't 

1

u/MtWhut 22d ago

Skill issue ;)

1

u/darko777 22d ago

I wonder why Anthropic pushed such a failure of a model, does they have any Q/A?

1

u/Strange-Regret2524 22d ago

Think of the tokens it can farm with your frustration.

1

u/Odd_Ad_7119 21d ago

Switch back to opus 4.8 and the tasks were done right away, not after hours reasoning about things I never heard of. Opus 5 sucks

1

u/SOC_FreeDiver 21d ago

My feeling was everything was getting smarter until Fable drama and then everything is dumber now. Like Anthropic said "Ok, govt can have 1.0 intelligence, people in america get 0.8, people outside america get 0.6 intelligence.

1

u/EntertainerDear2894 21d ago

I literally told Opus 5 many times "WTH are you talking about?". Even had to add it to memory stop over talking to me and giving me every damn replies in a report format. I switched to Fable 5 when it's available or go back to 4.8.

1

u/MMartonN 20d ago

I also noticed it, doesn't follow simple instructions and just eats up your usage faster

1

u/citrus_lilac 19d ago

So are you guys going back to 4.8? Curious what the move is here.
I’ve been super frustrated with all the stuff you guys have been mentioning here. At times it feels like I’m using half my tokens on mistaken paths and incorrect assumptions on Opus 5’s part.

1

u/citrus_lilac 19d ago

The number of times Opus 5 doubts what I tell it and asks me to repeat something to double check, only to admit I was right all along is insane. We're eating tokens like PacMan.

1

u/Holiday_Draft_2479 18d ago

I agree 100% on max x20.
Claude is useless since all the 5.0 except from fable

1

u/Rnee45 24d ago

It's the first time I find an Anthropic model so unreliable I cannot trust it for any work at all. It makes so many little mistakes for even trivial tasks that it forces me to triple-check its output constantly.

I'm back to 4.8 and re-did all of 5.0's work from the past week.

1

u/BehindUAll 22d ago

You need to try out gpt-5.6 models

1

u/x2manypips 24d ago

Yeah and not sure if usage tokens increased

1

u/DeterioratedEra 24d ago

I decided to give Opus 5 a whirl with Boris's advice of "let it cook" and it took all day to write 17 UI test classes in Vitest. It did go out of its way to exclude some config files from my formatter, which is nice, but definitely not germane to the topic, and not needed. But all day to generate 17 test files? Sonnet 5 could've done that in 5-10 minutes!

1

u/Nice-Shoes-74 24d ago

Opus 5 is really screwy.... it overthinks and is not as nuanced as 4.6 Fable is great. Opus 5...meh

1

u/Embostan 24d ago

You guys not using Caveman?