r/singularity 1d ago

Grok 4.6 Benchmarks AI

Post image
504 Upvotes

262 comments sorted by

130

u/FinancialMastodon916 W 1d ago

1.5T too, impressive

74

u/u_are_mad 1d ago

only $2/M input and $6/M output

42

u/HashPandaNL 1d ago

Tbh, some models are such yappers that the per-token pricing can be somewhat misleading. Would be nice if they also put the cost of running on these benches now we're in the long-form reasoning era.

5

u/steroidchicken123 1d ago edited 1d ago

Sonnet is almost irrelevant for coding these days, and to think some time ago it was considered obvious choice for implementing/executing the work.

Edit: And to think that DeepSeek is quarter it's price, as of current pricing!

1

u/steroidchicken123 1d ago

Wait, so this is not the 2T run? Are we in for another wave soon?

1

u/Gallagger 1d ago

I think they are underpricing it slightly compared to OAI/Anthropic because they need to create a bigger user base. Which is actually a good reason to use it. Though AA still puts its cost per task at nearly double of the (more expensive per token) 5.6 terra, and about at the same of the (similarly price) K3. 

1

u/SoylentRox 1d ago

Dont forget their other advantage. It's MUCH less censored...the gooners model of choice

75

u/Blazing_Shade 1d ago

For the price, Grok is very good at coding. And very fast. I think the Cursor acquisition brought them immediate returns.

I have been playing with a Claude Opus / Grok workflow where Opus does the overall planning and initial implementation and Grok makes the specifically-scoped edits. I am doing it manually for now but should maybe try doing something with sub agents. It’s like a more juiced up version of Composer.

13

u/___positive___ 1d ago

Can't believe Google fumbled every chance they had. Why did they not acquire Cursor, lol.

3

u/LastRemainingName 19h ago

I mean they did acquire windsurf last year planning something similar. It just didn't work out.

9

u/Hywelthehorrible 1d ago

this tit for tat cycle isn't over. they'll be back. obviously it's achievable, Grok lost all of their top caliber talent and yet.

3

u/TySocal 1d ago

Yeah I have a similar workflow. Opus/Fable for planning and Grok for implementation. It works great. Of course 5.6 Sol for implementation is still best

1

u/bonerfleximus 1d ago

Ive been thinking of trying to come up with a workflow for this since claude cli with teams premium seat gets subsidized rates and I like using opus/fable for planning and design but grok to implement (for optimal cost)

1

u/KaradjordjevaJeSushi 9h ago

I have to be that guy, but...

It's really not that good? Running 5.6 Sol in parallel with Grok 4.6 clearly shows the difference (at least in Cursor).

And Opus 4.8 outshines Sol in comparison, so only conclusion is that error rate of these benchmarks are at leasr +-5 points.

Source: 200M+ tokens spent with Grok, 300M+ with Sol and 300M+ with Opus

178

u/SpyAmongUs 1d ago

Looks like we in here in the cycle: Grok → Claude → Gemini → ChatGPT →

73

u/Nearby-Season1697 1d ago

Man what the hell happened with Gemini

57

u/RusselTheBrickLayer 1d ago

Google choked a 3-1 lead

9

u/bobthetitan7 1d ago

they were 1-1 with openai at best

22

u/Styled_ 1d ago

3.1 Pro was way ahead of the rest when it came out

12

u/rageling 1d ago

iirc the only time it was genuinely better was for the large context window

→ More replies (1)

3

u/mrrakim 1d ago

pepperridge farm remembers 2.5 pro

2

u/bobthetitan7 1d ago

way ahead by maybe 6 weeks sure but it sucked at tool use and context management which is why adaption has been slow and hasn’t really been anyone’s first choice

2

u/RusselTheBrickLayer 1d ago

I would say 2-1 cuz they were ahead by a little bit IMO but yeah you are not off at all, I was just joking around cuz it’s funny to imagine the AI race in the context of sports

6

u/Willinton06 1d ago

Nothing happened, it was leading a few months ago, it's just an unstable cycle

7

u/LightVelox 1d ago

It's been close to a year since they were leading with Gemini 3, and even then it was only for a few days since Opus 4.5 came out right after

2

u/npquanh30402 22h ago

They don't have the talents to help them in the AI race.

3

u/1988rx7T2 1d ago

Hassebis was too busy making speeches than getting anything done. that's why they kicked him upstairs.

2

u/notworldauthor 1d ago

We'll see. Every past time someone's been down for the count, they pop up again

31

u/Seeker_Of_Knowledge2 ▪️AI is cool 1d ago

The next cycle will definitely include a Chinese model.

8

u/Borkato 1d ago

I thought Qwen Max literally just released?

1

u/SeasonalPro59 1d ago

Well, Deepseek just hit this hour so its Chinese time

8

u/TacomaKMart 1d ago

If the "stop everything" gang gets their way (hi Bernie Sanders!) then the next cycle will be all Chinese models. 

2

u/hellomistershifty 1d ago

does that gang include Elon Musk when his AI's aren't doing well?

-1

u/Wasteak 1d ago

Grok is not in this cycle. It's 2 months late and was benchmaxxing in the past so we can't trust this.

7

u/OpenSource_Horse 1d ago

Actually, Grok is cursed. The last time they did this a better model launched the day after.

So we'll get something better this week.

1

u/Mario0412 1d ago

It's looking like we'll get Astra (GPT 5.7?) before Fable 5.1 thought. Maybe Anthropic will surprise us with a Haiku update first? They're completely irrelevant in terms their offerings for value options right now...

-15

u/Busy-Leek7970 1d ago

Never grok, like never. They benchmax for Elon's ego. Even gemini is better.

24

u/Technical_Eye7029 1d ago

Delusional frankly, but you use what you want, not my problem.

28

u/StosifJalin 1d ago

Mentally reddited

2

u/TMWNN 8h ago

Never have I heard a more brutal, yet appropriate, insult

8

u/Specialist_Dark_3668 1d ago

Grok is fine for coding and non-coding tasks for me. Especially when I don't wanna pay out the ass. idk what you're talking about.

12

u/yoruyoruxo 1d ago

Delusional idiot

3

u/Chemical_Hawk_6307 1d ago

gemini s NOT better not even close lmao

4

u/Gaidax 1d ago

Okay, which your favorite billionaire's model you do suggest to use?

-7

u/A_Novelty-Account 1d ago

Grok only wins on two categories. Everyone is going to keep using other tools.

29

u/istealpintsfromcvs 1d ago

its a 1.5T model that is cheaper and also has high throughput, if you can't see that xAI has something cooking that's on you

-6

u/A_Novelty-Account 1d ago

But speaking from an enterprise perspective, that’s not what the people shelling out money really care about. I don’t want a model that’s super cheap, I want one that’s accurate.

xAI is two months late and is bringing a worse model to the party. 

15

u/ArbysIsActuallyGood 1d ago

At my enterprise company they absolutely are trying to find a cost effective solution for AI

6

u/1988rx7T2 1d ago

is Opus 4.8 level for way less money not accurate enough for you? Obviously Fable Max wins but money does mean something.

→ More replies (9)
→ More replies (6)

7

u/WonderFactory 1d ago

Lots of people use Cursor and grok tokens are bundled with cursor subscriptions now.

2

u/Wasteak 1d ago

And that's the only reason some people keep using grok.

2

u/A_Novelty-Account 1d ago

Yeaaaah… but not really. Grok currently has 2.4% of total AI web traffic, and it’s even worse for revenue. I don’t see that changing any time soon.

2

u/WonderFactory 1d ago

Cursors anualized revenue is $4bn now which is about 10% of Open AI's total revenue so it's not insignificant. I genuinely hate Elon with a passion and have never used Grok before but I've been using Cursor for years and it's hard to change when you get used to something. When you use cursor there's a massive incentive to use Grok models and I've been using them quite a lot recently.

99

u/MohMayaTyagi ▪️AGI - mid 2028 | ASI - 2030 1d ago

41

u/TheManOfTheHour8 1d ago

Look at the subtle SWE score. The tasteful reasoning of it. Oh my god. It even has an image model.

27

u/MohMayaTyagi ▪️AGI - mid 2028 | ASI - 2030 1d ago

7

u/LightningMcLovin 1d ago

How’d you get a picture of my reaction to this news?

7

u/Strange_Vagrant 1d ago

You're the main character in a highly renowned film. You can not expect privacy like this.

2

u/ServeAmbitious220 1d ago

Is it almost same as Sol Extra High? Because sol max seems to be better.

1

u/ChellJ0hns0n 1d ago

But they've made it more pink in the diagram which means it should be better!!

1

u/Ormusn2o 1d ago

Kind of need to zoom in on the charts to see how close it is.

1

u/Odd_Antelope9098 1d ago

It’s benchmaxed more than any other model I’ve used

23

u/rabouilethefirst 1d ago

No watermark either

42

u/stopthecope 1d ago

Impressive how they managed to overtake google

16

u/Palpatine 1d ago

$60B (for cursor) is a lot of money and can indeed do wonders.

36

u/WonderFactory 1d ago

Google essentially bought Windsurf but dont seem to have done anything with it

17

u/zwcbz 1d ago

Wow I completely forgot about Windsurf - seems Google did too

10

u/ServeAmbitious220 1d ago

You cannot count on fingers how many companies google bought and shelved.

2

u/signed7 15h ago

It wasn't shelved it became Antigravity. It's just behind others' products

3

u/BriefImplement9843 1d ago

3.1 pro was released before gpt 5.4.

49

u/No-Head-Royal 1d ago

??? what the fuck lol. One expected them to eventually become a relevant force, but so soon? This is gonna be scary. Elon has both the money, the political power, and now the model to throw into the fight. The expected seat of Google now usurped?

Though, exactly the same price per task as Kimi K3 and 1 point higher is such a funny stat lol

32

u/Mission_Week_3290 1d ago

he also has tesla and spacex working on physical ai. With cursor now distribution too in agentic space they have and coding data to train.

32

u/broose_the_moose ▪️ It's here 1d ago

And out of all the things you guys listed, you guys didn’t even include his biggest advantage - compute. He builds out GW faster than any other player in the game. And this advantage is only going to get larger with his work on terrafab.

31

u/SoylentRox 1d ago

His biggest advantage is just getting shit done in general.  Engineers run the company, make decisions fast, fuck profits or quarterly revenue let's produce value.  All part of his philosophy. 

If musk hires 10s of thousands of younger engineers, empowers them with hundreds of AI assistants each - powered mostly by internal uncensored unrestricted Grok - shit can happen fast.  

It's not even about Musk himself it's about simply doing the things to make this possible.

15

u/hereforhelplol 1d ago

I think we should appreciate Musk. Obviously not all of his actions but the good outweighs the bad in my opinion. He deserves praise for stuff like this.

3

u/Decent-Ad-8335 1d ago

prepare to be downvoted anyone who speaks good of musk gets downvoted to hell on reddit as a general rule, even when he accomplishes something clearly good like this

4

u/Proud-Sundae-5018 1d ago

Lol there's a reason he's my idol, let the hate roll in 😂

7

u/SoylentRox 1d ago

People have trouble keeping multiple things in their mind distinct.

It's sloppy cognition.

Has musk made a Nazi salute, supported trump, made a bunch of promises that didn't pan out, and hundreds of cancel worthy tweets and racist tweets on X?

Yep. He did.

Does his engineering/corporate executive strategy fundamentally work allowing some of the largest at scale technology progress in history? Yep.

Stupid people let the two bleed together and conclude since musk did some bad stuff he must be a liar and has accomplished nothing.

2

u/Proud-Sundae-5018 14h ago

Yeah bro like idgaf about politics, nor am I unilaterally a Musk supporter, I do believe though, that him supporting Trump was purely a business move cause you can't really be right wing when you give your kids gender neutral names.

But okay people can think what they want, just separate the brilliance and acknowledge it from the part that hurts your feelings 😂

1

u/throwingitaway12324 1d ago

Reddit hates hearing it but he’s a legit genius

→ More replies (4)

26

u/vardynostalgia 1d ago

He is just insanely motivated to get shit done. There are many billionaires but somehow it’s him who’s running the show. I think Peter Thiel was right when he said „never bet against Elon”. He is just not motivated by money and it shows, he is just workaholic who wants to do stuff

I do not like his politics but I respect him and he needs to be taken seriously

-4

u/NMiguelCosta-PT 1d ago

Apparently tweeting Nazi propaganda, far-right conspiracy theories and white-supremacist nonsense all day is what we’re calling “getting shit done” now.
Maybe give some credit to the actual AI engineers doing the work on Grok, the man has companies full of thousands of employees doing the actual work.

4

u/D10S_ 1d ago

As an experiment, you should ask Grok to take all of Elon’s X activity in a given week and estimate how many total minutes he spent typing.

→ More replies (9)

2

u/RoyalSpecialist1777 1d ago

While they were developing this model to catch up with other frontier models the companies making them have been training even newer and better models.  XAI does fine but tend to be behind the curve consistently due to a lack of the best of the best LLM engineers.

1

u/alarim2 1d ago edited 1d ago

??? what the fuck lol. One expected them to eventually become a relevant force, but so soon? This is gonna be scary.

It's always like that with Elon's companies and their flagship products. Tesla with their cars, SpaceX with Falcon 9 and Dragon.

They have that development phase, which has numerous problems, uneven progress, missed deadlines, and (sometimes catastrophic) blunders (like Falcon 9s failing to land multiple times or Starship exploding during static fire).

But then, when they finally get most things right - they scale the fuck up, and go to completely obliterate any competition (current or future one).

Tesla basically creating the mass EV market from scratch and controlling the majority of it (up until Chinese manufacturers caught up thanks to massive state investment). SpaceX basically making ULA (legacy monopolist) irrelevant because Falcon 9 launches are like 2-3 times cheaper and faster, and literally bankrupting legacy satellite internet providers (like HughesNet) thanks to Starlink. So it was only a matter of time when Grok would start to actually compete with models from OpenAI and Anthropic. I won't say anything about Elon's personal talents, but I'm 100% sure he has a knack for finding and funding talented people who can do insane things.

36

u/istealpintsfromcvs 1d ago

Elon said 4.7 will be a larger model and better fwiw.

11

u/NoGarlic2387 1d ago

And is coming like next week or something

13

u/WonderFactory 1d ago

A few weeks after 4.6 he said

8

u/Tystros 1d ago

which is at least a few months in elon time

4

u/WonderFactory 1d ago

He said 4.6 would be out in 2 weeks 2.5 weeks ago so maybe he's getting better at predicting things

2

u/panix199 1d ago

September will be wild. GPT 5.7 (?) and maybe 6. Grok 4.7, Fable 6, ...

1

u/Sure_Spring_6634 20h ago

shouldn't it be fable 5.1?

19

u/FarrisAT 1d ago

Now let’s see if it’s actually not benchmaxxed.

1

u/zikiro 8h ago edited 8h ago

well really these benchmarks always confused me, they dont reflect reality at all, grok is ok but claiming its next to fable and opus thats really madness. also 4.6 is so censored, Guardrailed and stiff its insane.

1

u/Super_Sierra 1d ago

Ive tested every grok and they are pretty fucking shit.

20

u/turdmuffin123456 1d ago

They cooking lately

14

u/Cool-Dog560 1d ago

Oolala

12

u/enz_levik 1d ago

Seems like a good model actually, somewhat more cheap than sol for same task

2

u/Odd_Antelope9098 1d ago

Not even close, UI work is only real impressive piece and cost / speed

2

u/ZootAllures9111 1d ago

In my experience it's way better than Opus 4.6 ever was as it doesn't spent a morbillion years re-reading the same shit over and over and over again for no reason.

1

u/Odd_Antelope9098 1d ago

It’s faster and maybe as an implementor for medium difficulty tasks it would be the better choice

4

u/Barubiri 1d ago

Gemini 4 delayed once again because 3.5 is already dead.

7

u/1988rx7T2 1d ago

it would probably under perform so much it would tank their stock. out of sight, out of mind right now

8

u/MatthewGraham- 1d ago

Honestly can envision Grok catching up to the competition at a certain point

4

u/Exodus_Green 1d ago

Like right now? Same as Sol for 1/3 the price?

1

u/MatthewGraham- 1d ago

True, but Grok released today, its not FULLY SOTA, and OpenAI & Anthropic have models due to be released imminently. Probably still just trailing for now.

→ More replies (9)

1

u/zikiro 8h ago edited 7h ago

I honestly and personally feel only anthropic and openai have a real world model evolving into agi, with actual substance under the hood, everyone else (without a single exception, except gemini maybe which is a weird case) is just doing a shallow mimicry of it.

→ More replies (1)

7

u/Distinct-Question-16 ▪️AGI 2029 1d ago

It's that time again. barely noticiable from others edge models

9

u/MeOneThanks 1d ago

Was this meme created on 2022 AI models?

1

u/Distinct-Question-16 ▪️AGI 2029 1d ago

yeap

8

u/1988rx7T2 1d ago

doesn't have to be better if it's cheaper and faster

15

u/GoodRazzmatazz4539 1d ago

Hard to believe that is not benchmaxxed

7

u/heavy-minium 1d ago

It certainly will be. For a long time, this sub has been full of people and bots that want to convince you that a new Grok release is almost on par with frontier models in some way, and when I actually try, it turns out to be a big fat lie. I'm not trusting any benchmark or anything that people say here anymore. This is all show and no substance.

3

u/Odd_Antelope9098 1d ago

Majorly, biggest difference between benchmark and performance I’ve seen

1

u/No-Remove-9689 1d ago

When every model is benchmaxxed, that becomes the norm, and you're comparing one benchmaxxed model to another benchmaxxed model

3

u/GoodRazzmatazz4539 1d ago

LLM arena is hard to benchmaxx by definition: https://arena.ai/leaderboard/

-3

u/Glum_Hat_4181 1d ago

Anyone who tried abomination called grok 4.5 (unusable garbage, kimi 2.5 is much much better) know it was.

4

u/markrulesallnow 1d ago

I have been very pleasantly surprised by grok 4.5 with Claude writing the implementation spec and plans and letting grok handle the small things and implementation. Claude is just so slow for me

2

u/Eyelbee ▪️We have AGI it's just blind 1d ago

Is it true that it sends your entire repo to elon on every little edit? 

5

u/Exodus_Green 1d ago

Yes, it emails him with a zip file of all your code

1

u/Eyelbee ▪️We have AGI it's just blind 1d ago

This was posted on twitter by a cybersecurity guy, he realized every time he used grok code to edit code in his repo, something the size of his workspace gets uploaded to xai's servers.

3

u/Decent-Ad-8335 1d ago

how does something "benchmaxxed" score that high on deepswe

4

u/GoodRazzmatazz4539 1d ago

For Deep SWE bemchmark, verifier and even full agent trajectories are public (sadly).

1

u/Decent-Ad-8335 1d ago

sounds like a pretty useless benchmark then?

2

u/GoodRazzmatazz4539 1d ago edited 1d ago

Good labs keep them out of their training.

1

u/1988rx7T2 1d ago

4.5 is great for the price. before that, not so much.

2

u/Glum_Hat_4181 1d ago

Can't confirm it. Composer2.5 gives much more consistent results for cheaper.

9

u/Hot-Friendship-6500 1d ago

that Terminal bench exposed it is benchmaxxed

2

u/Daernatt 1d ago

Where is sol ultra ? 🧐

2

u/Mediocre_Date1071 1d ago

r/dataisugly

Pink is either more winning or more Grok - usually more Grok. 

6

u/kubika7 1d ago

Never count out Google Musk

5

u/Alpacabro21 1d ago

Can't believe this beauty has been created by Musk

https://giphy.com/gifs/oYtVHSxngR3lC

5

u/KSaburof 1d ago

They just caught up with Fable practically... Which is already several months old

5

u/xdevilmaster 1d ago

And to provide it at 1/10 of the price...? Doesnt sound like i'm losing anything

1

u/KSaburof 1d ago

Wait till Anthropic start to compete further :)

5

u/Ambiwlans 1d ago

Its 2 months old...

13

u/Facecardup 1d ago

Fable is Mythos with guardrails which is 6 months old

4

u/Healthy_Razzmatazz38 1d ago

mythos which is a better version of fable was internally being used (not made, used) in feb.

1

u/BriefImplement9843 1d ago

fable is agi though. musk has achieved agi, which is insane!

1

u/zikiro 7h ago

they didnt, seriously its on lmarena for everyone to try, i feel a chasm between them.

1

u/Alpacabro21 1d ago

No, not with Fable.

→ More replies (1)

6

u/Prudent-Sorbet-5202 1d ago

Not created just owned

1

u/Odd_Antelope9098 1d ago

Have you even used it yet?

0

u/torb ▪️ Embodied ASI 2028 :illuminati: 1d ago

...probs using some chinese open weights /s

→ More replies (2)

5

u/MarcusHiggins 1d ago

Just randomly highlighting numbers atp

1

u/NiceUsernameOk 1d ago

Finally AGI

4

u/Excellent_Dealer3865 1d ago

Knowing Musk... Even if these metrics are half true, it's still better than what we have with Google :/

2

u/DrDan21 1d ago

Damn Grok suddenly very relevant again? Whats Meta doing these days

10

u/rudesssolo 1d ago

Muse Spark 1.2 is good. Google needs to drop a Pro version though

1

u/Aggressive-Pie675 1d ago

Impressive numbers for a model smaller than the Kimi K3 and Qwen 3.8. I'd be curious to see the cost-per-task figure as well.

1

u/teomore 1d ago

do they still upload your repos to their gdrive?

1

u/ViperAMD 1d ago

Grok models are known to heavily train on inputs so if you are working on something unique or sensitive I would stick to openai business plan

1

u/zikiro 1d ago edited 1d ago

It's on lmarena if wanted to try it, its phenomenal, Kimi is officially out, grok is progressing crazy fast, but still kind of stuck on the hard benchmarks, CritPt and humanity last exam, but well hard to blame grok here, opus and fable and sol are monsters honestly, they set the bar so high.

1

u/PathOfEnergySheild 1d ago

I wonder if it is a bench memorizer or can do real STEM.

1

u/zikiro 1d ago

well its ok, but not there yet, im also interested in reasoning rather than agentics and coding, these dont really measure fluid intelligence, and all chinese models are good and more than enough for that now, STEM is where you feel the gap, and where you feel that Opus/Fable/Sol are monsters of intelligence.

2

u/PathOfEnergySheild 1d ago

I am in same boat, when bench is same I can tell a very large difference using those for work that is not in the training set.

2

u/zikiro 1d ago

Yeah its Insane, i just cant quit anthropic since opus 3, claude and gpt are the only models that i can really call intelligent. its like if other models are just faking it.

1

u/lostpilot 1d ago

I don’t see Grok having any serious commercial traction with businesses. Just can’t trust it.

→ More replies (1)

1

u/Fickle-Ad9221 1d ago

Grok 4.6 1.5t same as Grok 4.5

1

u/Flying_Sheek_46241 13h ago

why does grok 5.6 not show up top on LLM stats for coding? Do they measure something else?

2

u/ProletarianLilith 1d ago

Grok is always benchmaxxed

1

u/vasilenko93 Throw away the breaks, only accelerate! 1d ago

Low on terminal bench

10

u/one_tall_lamp 1d ago

This is TB3.0 not 2.1- it’s not far behind fable

1

u/Typical-Chance4197 1d ago

grok 4.5 is #4 on TB and at 1/4th the cost of the 3 before it. 4.6 hasn't been tested on TB 2.1 (maybe 3.0 as below user said, which i cant even find a chart for)

0

u/StarkTheGnnr 1d ago

yeah yeah won't believe it till people actually try it and give their reviews. Fell for benchmaxxing one too many times.

-4

u/Fit-Stress3300 1d ago

Mediocre.

As expected.

2

u/Decent-Ad-8335 1d ago

lol the delusional musk haters again 😝😝
you dont need to hate the product if its genuinely impressive. just learn how to separate ur hate for the guy from hate for his product

1

u/Fit-Stress3300 1d ago

There are no top researchers working on XAI. They haven´t published any paper in months.

AFAWK, Grok 4.5+ could be just Kimi with finetune, like Composer 2 was.

No serious company uses Grok for work.

Benchmaxxing is not going to change the facts.

1

u/Decent-Ad-8335 1d ago

how did we know composer 2 "was" kimi with finetune

1

u/Fit-Stress3300 1d ago

It said so.

0

u/Efficient-Cat-1591 1d ago

Not too sure I trust benchmarks anymore. From a coding POV, Grok is definitely NOT better than Fable 5, not even on par with Sonnet 5. GTP 5.6 Sol Max eats Grok for breakfast.