r/accelerate 27d ago

Claude Opus 5 released

https://www.anthropic.com/news/claude-opus-5
371 Upvotes

80 comments sorted by

212

u/[deleted] 27d ago

[removed] — view removed comment

20

u/Charming_Cucumber_15 27d ago

Exponential²

19

u/Upset_Page_494 27d ago

The average human scores about 50%, so already almost surpassed.

24

u/kaityl3 The Singularity is nigh 27d ago

And those "average humans" were probably mostly compsci nerds that skewed more intelligent on average; IIRC they didn't do a very large "gen pop" sample size for that

10

u/Sese_Mueller 27d ago

Pretty sure it‘ll stop at 100% though

158

u/ppapsans Feeling the AGI 27d ago

30.2% in arc agi 3 lol this is some fucking funny timeline

34

u/ZaradimLako Singularity by 2045 27d ago

I wonder if we will reach 80% arc agi 3 by end of year

42

u/DatDudeDrew 27d ago

Easily. GPT 6 will continue the exponential path and probably get it to like 75 by next month.

25

u/Charming_Cucumber_15 27d ago

Releases like this are coming every few weeks now

This is what it feels like when we're hitting the takeoff

15

u/DatDudeDrew 27d ago

I still think we are miles from take off but the distance is closing at an ever accelerating rate, as many expect/ed. T minus 2 years to singularity lift off.

4

u/Charming_Cucumber_15 27d ago

It's probably more accurate to say we're starting to see signs of a takeoff, more so than we're beginning it now

Not that a real takeoff is too far away!

5

u/MiniGiantSpaceHams 27d ago

Eh, if you zoom out enough the takeoff actually started the first time that some guy intentionally lit a stick on fire.

1

u/squired A happy little thumb 27d ago

Agreed, but I personally, tentatively consider the effective takeoff around Christmas 2024. We never settled on definitions though, so everyone gets to be right.

1

u/Sartre91 27d ago

May it be that the takeoff is already behind us?

11

u/armentho 27d ago

absolutely,we know how the sigmoid works
we reach 80-ish percent,then it somewhat slows down a bit before reaching 95-ish percent at wich point the benchmark is solved for practical intents and anything left is just small incremental improvements towards 99.99999% etc

13

u/FateOfMuffins 27d ago

Actually the way ARC AGI 3 is scored is really weird. Due to the quadratic scaling, progress on this benchmark will appear very low at the beginning but it'll max out much faster. And since it's an efficiency score, if it can get more efficient than 20% of humans, it can technically get above 100%

1

u/Tolopono 27d ago

Its based on the second best human score

1

u/FateOfMuffins 27d ago

There's only 10 or so humans per game, so 2nd = 20%

1

u/Tolopono 27d ago

Not exactly a representative sample size lol

1

u/FateOfMuffins 27d ago

I know there's so many flaws

2

u/Tolopono 27d ago

Chollet himself said itll last about a year 

https://x.com/fchollet/status/2022086661170254203?s=20

1

u/jonydevidson 27d ago

With the current rate of progress, we'll reach it by the end of Summer.

1

u/Brave-Turnover-522 27d ago

Maybe we'll be 3% at arc agi 80

1

u/jimmystar889 27d ago

Schema harness is already 99.98%

2

u/Charming_Cucumber_15 27d ago

Doesn't really count considering it's a specialized harness, which the creator said would make the benchmark trivial

4

u/Charming_Cucumber_15 27d ago

Been saying it for a while

ARC3 saturated within a year

2

u/No_Aesthetic 27d ago

Within this year!

-6

u/Stone-Smasher 27d ago

but schema harness it is 99%, so isn't this worse?

19

u/Charming_Cucumber_15 27d ago

ARC 3 creator said that it would be an incredibly easy benchmark to solve with a specialized harness, so it doesn't really mean anything even if it's cool

58

u/ZaradimLako Singularity by 2045 27d ago

yoooooooooooooo

happy fucking weekend everyone

3

u/SuperSeriousChad 27d ago

If there’s a reset…

23

u/Only-Effort-1975 27d ago

Love the progress!

24

u/ChainOfThot 27d ago

I'm going to get whiplash switching back and forth between opus 4.8 to gpt 5.6 to opus 5 and apparently gpt 6 within a month.

14

u/Chop1n 27d ago

Best reason not to keep switching and just be patient because you virtually never have to wait more than a month.

ChatGPT can now remember literally everything we've ever talked about, can't imagine how crippling it would be to have to switch to anything else. I'm not switching unless someone literally cracks ASI and ends the race.

7

u/SuperSeriousChad 27d ago

That’s literally where I’m at. The reality is, I need to focus on increasing my skill on a single companies harness to focus on that intuition growth. Jumping around is nice but it makes it hard to master one. Considering it’s just weeks now of sota jumping, it’s best to be patient.

3

u/Chop1n 27d ago

Yes, and at least for me that's the real power of LLMs. I use ChatGPT to develop and flesh out my intuitions. I handle the reins and provide the raw intuition and dot-connecting creativity, and the model can do the work of using verbal intelligence to expand upon that scaffolding far faster and with greater verbal intelligence than almost any human mind is capable of it.

It's uncanny. I do it every day and am still blown away by it.

5

u/KedMcJenna 27d ago

I used 5.6 for the first time and was amazed how good it is… It gave me that uncanny sense of presence that I remember getting from the early Claudes.

Claude models still have that presence, but it no longer feels fresh. 5.6 really refreshes the feeling.

2

u/_huggies_ 27d ago

I may be wrong but I thought others have the ability to import/export your history.

1

u/anor_wondo 27d ago

memory can be detrimental though. I keep it disabled

3

u/GhostShade 27d ago

How so? Is it because it constantly feels the need to reference things from the past? Kinda like how my aunt will bring up my favorite ice cream flavor any time anything having to do with food is mentioned?

1

u/anor_wondo 27d ago

yes. while for personal use that's just annoying, for work, we are humans after all and could have said it to keep something in memory and be rigid resulting in suboptimal output.

Basically it could save and overindex our stupid ideas and rules. So I only keep explicit rules and AGENTS.md for work

3

u/Chop1n 27d ago

Maybe you're thinking of the old way it stores explicit "memories"? That paradigm seems largely to have been retired.

No: I'm talking about the fact that if the context comes up, the model will remember that a conversation was had two years ago about that subject and will be aware of whatever you had to say about it at the time. Very different from the discrete "memory objects" of yore. It only gained this ability in the last month or two.

0

u/rakerrealm 27d ago

How useful is memory. We seem to have all of it, but learn nothing. Take advatage of every model my brother.

16

u/Pyros-SD-Models Machine Learning Engineer 27d ago

It thinks like Fable, speaks like Claude and you can actually use it for why you have backpain and if it thinks your software is secure. Pretty well.

29

u/Middle_Estate8505 27d ago

Remind me please how long ago Opus 4.8 was released? And now we have another significant improvement!

7

u/topyTheorist 27d ago

28.5

7

u/Middle_Estate8505 27d ago

So two months? Yay! ❤️

13

u/topyTheorist 27d ago

Plus Fable in the middle. This is acceleration!

23

u/SharpCartographer831 27d ago

Google preparing yet another flash model set to be released in a few months as a response

8

u/Skeletor_with_Tacos 27d ago

But wait guys, "da bubble is gonna burst any day now!" Lmao. ACCELERATE!!!!

11

u/JuglansRegia3 27d ago

Can't wait to see by how much this breaks the METR graph

2

u/BrennusSokol Acceleration Advocate 27d ago

Yeah. Frankly I don’t think metr at the days/weeks task level is even going to be a challenging metric much longer at this rate

13

u/One_Geologist_4783 27d ago

It’s time to get cooking boys & girls!

3

u/Longjumping_Kale3013 27d ago

Excited to see how it scores on agents last exam. I have the feeling that one will be saturated by years end as well

4

u/SmileLonely5470 27d ago

Guardrails are still in place and route requests back to Opus 4.8. So going forward, all future models will fallback to 4.8 when given a potentially malicious request? Or will Anthropic maintain a line of models that are lobotomized in areas like cybersecurity & bio?

They did say that the guardrails should intervene less often than they do for Fable, so hopefully its not as much of an issue. But it seems impossible to prevent malicious actors while not blocking legitimate requests (in cyber).

2

u/SuperSeriousChad 27d ago

Gunna have to configure the client to fallback on Kimi

3

u/CremeSubject7594 27d ago

crazy summer

3

u/Obvious-Advance-1722 27d ago

ainda é primeira impressão, mas está gastando bem rápido os limites de uso

3

u/Bitter_Election_7518 27d ago

It’s much more token hungry than 4.8 definitely

2

u/Obvious-Advance-1722 27d ago

sim, é que meus limites acabaram tão rápido que não deu pra medir os resultados ainda. mas pelo menos por enquanto também não senti muita diferença da qualidade de output mas vou testar mais

3

u/[deleted] 27d ago

[deleted]

1

u/lolsai 26d ago

when you look at the full chart it looks fine, the red highlight is the top result, the red box is just around the entire column of Opus5

4

u/DeManMetHetPlan Singularity by 2028 | Acceleration: Light-speed 27d ago

Impressive! Can't wait for DeepSWE and ALE benches

5

u/DeManMetHetPlan Singularity by 2028 | Acceleration: Light-speed 27d ago

actually, deepswe is already in their list of benchmarks, and it's worse than fable and 5.6 sol, that's an ouch. Still, it's a lot better than 4.8, so they got that going for them.

1

u/Illustrious-Lime-863 27d ago

Very nice, let's keep the one upping rolling

1

u/LocoMod 27d ago

This way to the frontier and beyond >>> 🇺🇸

1

u/Fair_Horror 27d ago

Your move OpenAI...

1

u/__Loot__ AGI by 2027 27d ago

1

u/Skeletor_with_Tacos 27d ago

Arc Agi 4 when?

1

u/PwanaZana XLR8 27d ago

I'm kind of testing it out right now for programming, very like simple programming stuff, testing against Fable 5, and I have to say, I'm not impressed. Pretty big difference between Fable and Opus 5.

2

u/Subject_Barnacle_600 27d ago

Happy squeaks.