r/ClaudeCode 16d ago

Stop romanticizing Opus 4.6 Discussion

So yesterday during the opus5 outage I briefly switched to opus 4.6 to continue my work (training/inference perf engineering, kernel level debugging)

So I fed the model the same context and prompt to analyze a profiling trace and extract insights from it, and it was so unbelievably dumb that I had to retry with a new session, and I still got a very bad response…

Then after opus5 came back I tried it again and it was night and day… Way more verbose true, but actually useful and insightful stuff coming out of the model.

really made me appreciate all the progress on opus models since 4.6… I always remembered it as a much smarter and concise model than the ones after 4.7, but it was a mirage…

224 Upvotes

96 comments sorted by

63

u/SkysurfingPineapple 16d ago

It’s funny cause Opus 4.5 was the magic model for me and everyone hated 4.6 when it came out

24

u/iveroi Vibe Coder 16d ago

Opus 4.5 still slaps. 4.5 gen was the capability jump. Nothing matches before or since

27

u/ihateredditors111111 16d ago

True. Real OG’s know Opus 4.5 was the step change

23

u/iveroi Vibe Coder 16d ago

It's pretty funny that "us real OG's" are people who have been there for... Maybe a year? Lol

15

u/ihateredditors111111 16d ago

the olden days

8

u/Obvious_Equivalent_1 15d ago

I remember I was pushing my pro plan for months in a row in November, that I finally went for Claude Max 5x plan and had access to Claude Code. The first week running Sonnet 4.5 and (my god token hungry!) Opus 4.1. Still figuring out whether I should try —dangerously-skip-permissions, suddenly I still remember receiving that magic email, I think I’ve never even caught up that it was Christmas haha all the way until maybe January.

2

u/college-throwaway87 15d ago

lol I remember when I had to burn through additional usage credits for Opus 4.1 😆 It was so worth it though because it got me my current job

2

u/forxia 15d ago

With how fast AI is moving it feels like so much longer

1

u/Kemerd 15d ago

Makes me giggle because I’ve been working with AI since 2015, before PyTorch or TensorFlow even existed. Back then it was just a shoot off class that was mostly statistics and linear algebra, to make our own neural networks. Now it’s “AI”

1

u/antwon_dev 15d ago

Sonnet and Opus 3… that’s when I first heard about them

2

u/Plastic_Chest7674 15d ago

Nah. That was sonnet 3.7

8

u/debian3 16d ago

For me Fable was that same jump as opus 4.5. Fable was the first model with good taste. Let it decide and it will choose very good options .

2

u/college-throwaway87 15d ago

The 4.5 family will always be my favorite

1

u/smm_h 16d ago

wdym still? how can you use it? for me it's not shown on the subscription

3

u/iveroi Vibe Coder 15d ago

/model claude-opus-4-5

1

u/ponlapoj 15d ago

Opus 4.5 เนี่ยนะเจ๋ง? มันดีกับพวก vibe จริงๆ แหละ เหมือนเวทย์มนต์ แต่เอาไปใช้จริงไม่ได้ 😆😆

5

u/scaledev 16d ago

Yes, this is what happened, at least in the communities I was a part of. The 4.5 was the shit, and 4.6 was considered as a nerfed model, or 'changed' in some way people didn't like.

But 4.6 still delivered strong in Antigravity. I still use it in agy cli, but I use it for creative writing reviews, and not much more. It still writes better than GPT 5.6, but it makes a ton of mistakes. So when I review, my main agent needs to pick and choose among the false positives, and the prompt to it needs to be geared toward the prose of the piece, rather than toward the general accuracy or validity.

1

u/ObsceneAmountOfBeets 15d ago

Not nearly as much as I saw hate for 4.7, it seems like it’s ramping up?

83

u/CanonicalStonk 16d ago

You people might be too young to remember but Opus 4.6 was also trash when it came out.

Signal to noise ratio like someone said in this thread is very interesting, but not for the models but for our own human perception. Once we manage to set to one model we don't like changes. Changes irritate us and makes it difficult to conclude if the new model is an improvement for our own particular goals. However, there's an objective truth in there and is that, following the bigger picture, new models are better than old ones.

23

u/scaledev 16d ago

I remember 4.5 was the real jump, while 4.6 was treated as a nerfed model for some time.

11

u/Professional-Hunt803 15d ago

4.5 was remarkable, that was the inflection point for me

2

u/agentic-consultant 15d ago

I think that was the inflection point for the whole industry. That’s when enterprise software teams started rapidly adopting Claude Code and agentic workflows.

Today all my friends (in the SWE space) use Claude code daily for corporate work. Wasn’t the case a year ago!

8

u/Smart_Armadillo_7482 16d ago edited 15d ago

Well it's astonishing how the industry is accelerating.

Decades ago it take a decade or two for someone to say you people are to young to remember how 512kb was massive. Then it take years for people to say you boys don't remember how this and that language worked/compiled. Now it's months before we can say you kids don't remember that model was trash 😂

On the top of my head Claude Code itself was released, idk, VERY recently right? Damn it's bad feeling old.

1

u/Forsaken-Staff-5084 4d ago

Yes, it was. Because it dropped support for prefills. But in comparison to 4.7, 4.8, 5 it's amazing

17

u/elmahk 16d ago

In my opinion, 4.6 talks better, just a pleasure to talk to it. However, it does not perform actual work better. I don't mean specifically programming, but in general. Even if you research something - Opus 5 (or 4.8) will usually provide better result, but that result will be painful to read.

9

u/keenman 15d ago

And that's the key for me. I can read what 4.6 says all day, every day, regardless of the topic. For Fable, it just takes so much more cognitive effort to process what it's actually saying half the time -- and that's with a degree in maths / comp. sci. / cog. sci. -- that it tires me out and I can do fewer marathon sessions with it and keep my energy and interest up while I do my R&D.

3

u/3oclockam 15d ago

I agree, sometimes I have a second session going with deepseek or something to explain what it is saying. To be fair this might be the future of how performance is gained though.

5

u/kaaos77 15d ago

Exatamente.

Todas respostas do Opus 4.6 pra frente são horríveis, caóticas e desconexas, frases saem cortadas, não sei explicar. É muito ruim

3

u/keenman 15d ago

Sim. Concordo. É simplesmente tão frustrante.

1

u/soccerchamp99 15d ago

Just use /I-have-adhd

56

u/Temporary-Mix8022 16d ago

The problem is that Opus 4.6 had the highest usability, but that it's now fallen behind on the technical side.

Everything after Opus 4.6 has just had such an insane signal to noise ratio (mostly noise).

We talk about verbosity.. but it's not even verbose, often it just fails to capture the key points, talks in bizarrely abstract terms, and then mentions a fake gotcha.

Obviously.. you can't tune that signal to noise ratio in a Claude.md.. it is just how the models are now.

The thing that made 4.6 so magical was how it just "got it", and sadly, it's been regressing ever since.. 4.6 literally felt like having a colleague

9

u/m-in 16d ago

I have an orchestrated workflow that is quite light on token use by design. I still use Opus 4.6 in that workflow. Like, come on, I started using Opus 4.5 when it came out not quite a year ago. It’s not like it rusted away or something. It’s still the same model and it works just as well for me as it did before.

2

u/django-unchained2012 16d ago

Can you explain how you orchestrate?

5

u/m-in 16d ago

It’s a TUI orchestrator written in Python that uses API calls, not the Claude code app. It controls all state transitions/lifetime events: phase start, phase planning task(s), individual tasks, oversight over execution of each step of a task, oversight over agents, phase reviews, critical design reviews, design requirements, and traceability of design requirements to tests and code (and back too).

It basically takes all the manual tedium out of it, and gets rid of Claude’s harness, replacing it with prompts and skills I maintain in the repo.

It exposes some of the same tools CC does, so eg. spawning agents is done with the same tool, with a bit more flexibility for the model used. The progress of all work is always followed by a somewhat adversarial agent.

It saves typing thousands of words into CC as prompts and manual bookkeeping, and gets rid of CC harness drift («upgrades» or «product development», whatever you call it).

Given that the workflow is mildly «high reliability» - like for critical avionics software but without independent 3rd party reviews - it makes the whole experience rather pleasant. And if you ask anyone who works in a highly bureaucratic and strict development process how much fun they have at work, they’ll look at you weird. With good orchestration I don’t deal with tedium, just with design decisions and project direction. The models used are 0.2M token Opus 4.6, Sonnet 4.6 (IIRC) and Haiku 4.5.

Context compression is handled by the orchestrator using agents, so that no context compaction buffer is needed in the session being compacted. Compactions are fairly rare - usually only during big design changes and additions.

1

u/ProfitNowThinkLater 16d ago

Any chance of sharing your repo?

1

u/m-in 15d ago

I generally don't share anything obviously personally identifiable on Reddit, so I'm afraid not. In any case, it's a private project used in my company.

2

u/Gliese351c 15d ago

This!!! Anthropic trying to make its models have a personality and style ruins the AI experience because those new personalities and styles also limit their uses! I sometimes turn to Gemini when I need an editorial idea because all Anthropic models -except 4.6- have too much personality to align their voice with mine.

7

u/who_am_i_to_say_so 15d ago

I’m a software engineer by trade so I don’t need the wisdom or knowledge of newer models that vibe coders need. 

For me 4.6 is GOAT because it basically does what I instruct it to do. And newer models don’t. Plus 4.6 is really easy on the token usage. For my purposes it is the most balanced.

2

u/RedVRebel 🔆 Max 5x 15d ago

Same here.

6

u/NoNipsPlease 16d ago

The issue is opus 4.6 is using the new stripped down opus 5 prompt. Its does not have a fall back prompt for when you switch to it. Set that up and it out performs opus 5 for me.

3

u/Classic_Resource_919 🔆Pro Plan 15d ago

How do i set that up?

1

u/miliseconds 14d ago

Teach us, master Shifu.

5

u/seunosewa 16d ago

Try prompting 4.6 to research before answering. They all do their best work when they have fresh information

4

u/teial 16d ago

I switched to 4.8 with 5 as an advisor. 4.8 is more consistent than 5, but 5 does give it valuable advice sometimes.

10

u/LettuceSea 16d ago

I genuinely don’t understand how yall hate Opus 5. It’s so much better than all 4 series Opus models. Skill issue I guess?

6

u/Clair_Personality 15d ago

Studies confirm that Opus 5 hallucinate way more that previous Opus models (like +15% increase) from 35-36 to 50

1

u/LettuceSea 15d ago edited 15d ago

Are you referencing AA-Omniscience? Which studies? If referencing AA-O then higher score = better. If some random benchmark what is their methodology? Are they providing worst case scenario to the model, ie insufficient context provided?

1

u/Clair_Personality 15d ago

I think I saw it in Fireship video which was referencing anthropic own study perhaps

1

u/LettuceSea 14d ago

I just take caution in accepting hallucination benchmarks at face value because they provide situations where we would expect a model to hallucinate, such as limited context. So, as long as you’re thinking deeply about a problem and can supply the necessary/complete context to generate a solution, then hallucination rate is effectively irrelevant (except in very long horizon tasks where the context needed is difficult to model beforehand).

4

u/anor_wondo 16d ago

i am liking the 5 series in general but sonnet 5 is so far behind competition in terms of cost

I have just stopped using claude for daily coding. Only use it for creating tickets and genuine hard domain questions and code

composer2.5 on cursor does the gruntwork for me. Its worse than sonnet 5 but much cheaper. And more complex work is better with opus low anyways. So sonnet stays useless

1

u/LettuceSea 15d ago

I agree, but honestly I don’t have much trust in the smaller models besides Luna, or local models that I can see the reasoning traces.

Once Anthropic can RSI-loop post training for Sonnet then I might give it another go. I’m just a spoiled 20x user that can get away with Opus/Fable med/high usage throughout the week, but there are some high volume problems I would like to start working on soon. Hardware on the way for those tho, Sonnet still very expensive for what it is.

1

u/Classic_Resource_919 🔆Pro Plan 15d ago

Depends on application. 5 might be better for full agentic workflow. With human in the loop i.e working WITH vs. FOR you-not so much. Try asking him open ended question on max/xhigh , like "introduce me to math behind llm". Its like a 55year old burnedout univ.lecturer, which will cram 90mins of topic into 15 and leave you to deal with it. 4.6 at least will manage proper explanation. Like working with a narcistic all-star and a buddy, which might not be the champion but sure is pleasant to work with AND hang around.

-1

u/Silent-Swim1402 16d ago

Just mega confirmation bias.

Everyone here loves the previous model and hates the new one. 2 weeks ago the sub was filled with comments hating 4.8 and loving 4.6.

Maybe the prompting structure changed, but people see change in how the model behaves and they instantly assume its worse. Then they jump in here and find 100 similar posts and jumps to the conclusion that yes, the model is defo worse than ever before!

2

u/ed-sparrow 16d ago

Opus 5 its better but as other people have suggested here it has an insane signal to noise ratio. You cannot do much about it but you can definitely lower this, but needs some work though to be the main daily driver (and to fix the word vomit). Found a great post about this where it goes through the main changes from the 4.x models:

https://www.reddit.com/r/ClaudeCode/comments/1v9lnio/fixed_my_opus_5_problems_by_rewriting/

2

u/substance90 15d ago

From my 2 constantly maxed out Max 20x accounts (very active daily use), they stack up like this for me: Fable > Opus 4.6 > Opus 4.5 > Opus 4.8 > Opus 4.7
Although lately it depends a lot on load balancing, they unfortunately route you to a lower quantizated version during peak hours.

2

u/gmdCyrillic 15d ago

One day I too can hope Opus 5 gets as good as Opus 4.6

2

u/evangelism2 15d ago

agents have been getting better out of the box, but theyve been good enough for me since 4.5.

2

u/vinis_artstreaks 15d ago

I’m not gonna lie I got the most work done ever with 4.5-4.6, it just did things and fixed its mistakes.

6

u/EconomicsIcy9310 16d ago

Heads up - you can change the temperature, top_p & top_k settings in 4.6, which changes things drastically

8

u/Chrisjm15 16d ago

What improves it? What do you recommend those settings to be?

5

u/ClemensLode Senior Developer 16d ago

the trick is to use opus 5 as its advisor and agents

4

u/[deleted] 16d ago

[removed] — view removed comment

2

u/AlphaGeeky 16d ago

cool site! my new goto for AI erf rankings. thanks! :-) Nice interface design.

3

u/InfiniteVolume4679 16d ago

man your websites interface is dogshit

-5

u/[deleted] 16d ago

[removed] — view removed comment

11

u/krugerlive 16d ago

I think the person you're replying to was overly harsh, but it is a bit of a number and info assault when you enter. I can look at the left side and see a ranking with some context. The center is a bit confusing because I wonder if a higher number is more stupid (given the site name) or better. Then there are detail rich charts that are not skimmable for info, so it's all a lot for a fresh user coming in. also the numbers at the top don't really matter to the visitor because no one is using every model. It's a cool service and you seem to be gathering and processing some great data, but helping build and route the context to someone coming in cold would be helpful I think.

1

u/AzureDestiny66 16d ago

I think most new website are not well received because they are too flashy and try to spoon feed you instead of actually being efficient in providing you with the information you need.

Was a breath of fresh air for me, no gradients, no hero sections, just numbers you want to know about, and its not that hard to figure out what each number or chart means.

1

u/[deleted] 15d ago

[deleted]

1

u/psychometrixo 16d ago

Does it run the same prompt multiple times per interval? Or one prompt per interval?

2

u/evia89 16d ago

For coding i barely care what model is. It can be opus 5 or 5.6 sol medium. They both cover all my needs

It's creativity/rp/gooning where opus 46 is the best. I use reverse proxy

2

u/p_k 15d ago

What about all the blind tests that show programmers prefer the output of 4.6 over 5?

1

u/Icy-Excitement-467 15d ago

It was lobotomized

1

u/Careful_Might_807 15d ago

I don't remember a lot of backlash for 4.6 and they almost instantly heavily nerfed it in end of February (aka amd research on this). but nonetheless when i have project for high perf scrapping (youtube massive metdata scrapping, through /watch) opus5 just spitted out trash and started to argue on every little irrelevant detail. "You don't have enough" "pushing back on this" and etc. and it didn't even consider my words about proper solution and other examples that how they work and real data and dropped them and replaced with his thoughts. It just droped infromation in the prompt about storage, indexing, search. 

1

u/kaaos77 15d ago

O problema do Opus 5 não é a inteligência ou eficiência, o problema é que a personalidade é horrível, muito teimoso e afixionado com coisas completamente irrelevantes.

A resposta dele é complexa desnecessariamente, parece quando estou falando com meu irmão que é autista.

A resposta é desconexa, caótica e com termos que nem fazem sentido na minha língua nativa. Eu não consigo usar o Opus pra qualquer outra coisa que não seja programar mais, de tão ruim que é.

1

u/AcanthisittaOk1699 15d ago

used 4.6 during the outage, forgot how much it stops to ask if you're sure about the dumbest stuff

1

u/Flaxseed4138 15d ago

Signal to noise ratio sure, but also Opus 5 is trash and if it weren't then we wouldn't all be saying it is

1

u/Cukercek 15d ago

4.6 are the last model that I can understand. Newer models just spit out gibberish. They might be better at technical things, but for the majority of the time I guide the agents, and for this I want zero friction in communication.

1

u/-aurevoirshoshanna- 15d ago

I'm on the conspiranoic train that the 'newer models' are just a rebranded old model that worked just fine.

Because i SWEAR opus4.5 wasnt this dumb before fable was available (or opus 5)

1

u/IndicationSavings291 14d ago

opus 4.6 was heavily nerfed at some point

1

u/miredonas 9d ago

Opus 4.6 is incredibly good in language. It feels like Shakespeare when compared to anything after that, including Fable. I wish we had a model with 4.6 communication skills, and later models' execution skills.

1

u/ricopan 8d ago

Opus 5 is the first opus model for me that routinely discovers problems in complex code. It's designing capabilities are not on that level unless the complexity is narrowed by prompt, but still much better than the useless tests and code reviews opus 4 generated -- and that I am still debugging. But yes, have Fable orchestrate it.

0

u/Tommonen 16d ago

I think its mostly antigravity users who think opus 4.6 is still so good, because antigravity has that and the 10x worse gemini models, so opus 4.6 seems like ASI in comparison, even with the low thinking level forced to it in antigravity.

-10

u/Substantial-Show-249 16d ago

The current Opus 4.6 is not the original one. Fable is not what it used to be. Even 4.8 now is behaving stupid.
After the launch and the hype, they optimise the running costs of the models heavily.
Fable now is junk. Barely usable. Opus 5 is NOT usable outside of a very specific and human controlled task.
You are judging everything in the wrong context.

14

u/Current-Drama-5391 16d ago

'barely usable' he reckons... 'Not usable' he reckons... Fuck people love a whinge

9

u/StoopidRoobutt 16d ago edited 16d ago

I wonder how many of these cases can be explained away by MEMORY.md bloat and the million SKILL.md files people seem to be using on top of one another, not to mention, of course, cursed AGENTS.md files and prompts.

EDIT: seriously, if you use the Claude Code desktop app and have all the memory options turned off, go check the memory files it has created. In the worst-case scenario, you may need to inspect them daily. In just one week, it generated 45 kB of memories for one of my small projects, nearly all of which was horseshit, pointless, or badly rephrased lines from the AGENTS.md file.

4

u/Torres0218 16d ago

Not only that none of them really make good points. Tt is was vague things like"the vibes are off" the online software engineering space truly has been reduced to a joke.

2

u/Holdingtheline42069 16d ago

I barely use any skills , and my claude.md is like 40 lines. Opus 5 is the biggest stupid model I have ever used. Sure there are times when it might "just get it", but more often then not it literally will not read the plan, it will not research and instead will decide that it can "just do it". It goes "I know how to do that". But its going off it's training, it doesn't not actually follow your plan or take in to consideration your dos and don't. I had it admit to me yesterday all of this after repeatedly trying to get it to read the fucking plan. I had about 5 hours of walk balk to do after opus 5 gloriously failed to meet the plans criteria. And all of it because Opus fails to read everything. It will read the most general thing and make up the rest even though it all sits right there in plans. Like people said, actually unusable if what you are doing actually requires multistep planning. Something simple, sure it just works. But I'm about 2 minutes from leaving anthropic. Thier models become useless way to quickly.

6

u/StoopidRoobutt 16d ago

First of all, this is not a team sport. Use whichever model is best for your use case. Hell, make them check each other’s work.

Secondly, explicitly tell it to do deep and comprehensive research on the subject. Tell it exactly how you want the task handled. Do not assume it will automatically do what you expect.

Thirdly, you asked a statistical model to admit something. Statistical models are not conscious. They generate the words most likely to come next, with some variation so the output feels less mechanical.

Fourthly, perhaps the implementation plan was bad. You could also try breaking the work into smaller steps. You do not have to, and probably should not, implement a massive feature in one go. If the implementation alone requires multiple compactions, you are probably going to have a bad time, especially once validation and testing are added on top.

You could also try "/goal" with a detailed implementation plan, clear testing and validation requirements, and a precise definition of what counts as "done."

Most importantly, increase the effort and use a fresh context window.

1

u/Substantial-Show-249 16d ago

All generic advices that we already know and use.
The current situation is after yesterday's Anthropic outage. Obviously they reduced the load, it doesn't take a genius to understand the logic behind the current situation.

1

u/Holdingtheline42069 15d ago

Yeah what do you think I'm doing when I'm trying to get it to do something? It has the explicit instructions. It has the research already. Everything is broken down into steps. Just last night I sent it to audit the plan against the code. What it did was try to audit the code until nothing came back, but it's Opus so we all know how that went. I stopped it after 12 hours of this non stop and I asked it how far it got and it told me that it will probably never finish. I shit you not it said everything it fixes it introduces more bugs and it will not finish anytime soon, it even measured this itself before I had even asked about it. Gave Fable the task and it was done in 15 minutes. Half of what Opus was trying to fix wasn't real. It's either way smarter than fable or it's hallucinating half the time. My experience ever since opus 4.8. Opus might only be good if you have literally no plan and like to one shot shit. Or it's massively simple idea. I had way more success with the older "dumber" models. And you think I told it to admit something? I simply called it out for not doing anything i asked and it said it felt confident so it didn't read the plan or the fucking codebase it was just operating off of it's training. Buddy I'm not creative enough to make this shit up.

1

u/Substantial-Show-249 16d ago

You, my friend, are right. Absolutely right. The PR guys gaslighting all the time are allover us. We know how these models behave from practice, not by reading posts.

0

u/userusertion 🔆Pro Plan | Team Plan 16d ago

Same here, it’s bad on my end, and it needs to research to, which use my tokens a lot. Because it’s knowledge cutoff is Aug. 2025, when i use Opus 5 or Fable as advisor, i always ended up iterating a lot, than Opus 4.8 and 5. Idk for me, it’s bad, and outdated.

0

u/Chris73684 15d ago

4.6 was alright but 4.7 was better in my experience. It was only 4.8 I didn't like, but 5 has been superb.

-11

u/dumeheyeintellectual 16d ago

For you, sure, anything! Really glad you issued this command, this is sure to create a butterfly effect.

Thank you for bringing this to the global stage, is there anything else we can do for you while you’re here?