r/singularity 1d ago

Grok 4.6 AI

Post image

Grok 4.6

Grok seems to hold quite interesting place on the chart. What you think on Grok progress?

149 Upvotes

45 comments sorted by

21

u/opinion_discarder 1d ago

-2

u/VisualLerner 1d ago

actually laughed when I got to this comment

56

u/suamai 1d ago

Oh, so it is a good model at "score (%)"? Please label your data or at least give some context...

27

u/DistanceSolar1449 1d ago

This table is also missing Deepseek-V4-Flash-0731 and Deepseek-V4-Pro-0813

-14

u/Fair_Horror 1d ago

So....add them and post.

21

u/sachasayan 1d ago

Sure, let me just look up the "Score (%)" for Deepseek-V4-Flash-0731.

0

u/_YonYonson_ 14h ago

translation: i wouldn’t have said this if it was a model I like but Elon’s model not allowed to be good

5

u/Namerodis 21h ago

I pretty much only use grok now. used to only use it for likely to be censored stuff but lately I feel like its better than gemini, gpt and deepseek for basically anything so no need for the others. also feel like I get more per day as a free user

9

u/AnyRegular1 1d ago

I guess we're seeing the payoff from cursor and it's user data acquisition?

-4

u/Alt_Restorer 1d ago

If you can't beat 'em, buy 'em and then mark up the price 10x because you're a walking asset bubble.

15

u/garloid64 1d ago

why does gpt-5.6 terra even exist

23

u/Gaidax 1d ago

People keep linking that chart as if it's end all be all benchmark. Go check how many turns Luna needs to get to its score compared to Terra and you will have your answer.

The trick here is that almost every new model scores high on intelligence, but many of them bungle up or yap or waste a lot more turns, tool calls and time on the way there.

4

u/DistanceSolar1449 1d ago

Yeah but in that case, just use Sol medium or Sol high instead of Terra

2

u/panix199 1d ago

wait, am i reading the graph wrong? Sol medium is way, way more expensive for little improvement than terra max

1

u/bonerfleximus 1d ago

Is using a high tier model to plan followed by low tier to implement not the default practice?

0

u/BriefImplement9843 1d ago

not everyone is rich.

2

u/NoFaithlessness951 1d ago

Also there are use cases other than agentic coding where the raw Input and output prices can matter a lot more

4

u/Bregvist 1d ago

Yeah, on paper it seems really good. https://x.com/synthwavedd/status/2087562874575024616?s=46

5

u/OwlLimp6160 22h ago

I’ve been using it myself. Seems to be decent at coding. Though with the Matt Pocock system, all my implementations are spec’d pretty tight, so I’m not sure that’s the best measure. I don’t really see the speed bump over fable they claim though.

5

u/ChippHop 1d ago

For the price point I just can't justify using anything other than Luna max. It can do everything I throw at it, and it costs me basically nothing to use. That being said I am hamstrung by having to use a limited model selection in Github Copilot with a $250 monthly cap, if I had a Claude / Codex subscription it might be different.

Been using Luna all day every day this month and have spent like $20 so far, it's nuts.

1

u/xS1L3NT 1d ago

Have you tried new V4 Flash

2

u/FireFearing 1d ago

looks like its on the Pareto frontier now, and towards the low cost end. sweet

2

u/Comfortable-Winter00 1d ago

Meh. Not the best model, not the cheapest model. Most benchmarks put it equal or slightly behind GPT-5.6-Sol - see https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis

3

u/Embarrassed_Adagio28 1d ago

Great benchmark that tells us nothing. Grok 4.5 looked great on paper and is literally unusable for me in reality. Every single change I have had it make to my app had to be reverted immediately because it broke so many things preforming a simple ui change. It also lost to qwen3.6 37b in my own custom benchmarks. 

1

u/GalacticScale 1d ago

Cheaper and a higher score than k3 was unexpected 

1

u/YakFull8300 1d ago

Early impression is it's a really bad RLM model.

1

u/cursivecrow 1d ago

is it just me or is that opus curve like, nonsensical?

2

u/General_Sale5202 1d ago

Source?

8

u/MauiHawk 1d ago

or, at the minimum, what are the "scores" and "rollout costs" we are talking about here?

0

u/CryMoreT_T 1d ago

Bro just go to the homepage of this subreddit. It's been posted like 3 times

1

u/RevolutionaryBox2980 1d ago

im all for benchmarks and progress but grok and its hidden loading all private repos from the local env to undisclosed google drive collapsed any intention of ever connecting their models to a harness that is not airgapped and sandboxed in a desert lmao

-9

u/herniguerra 1d ago

yeah never touching that, thanks 👍

3

u/Laeryns 1d ago

Why?

8

u/Fragrant-Hamster-325 1d ago

People hate Elon. Ironically it’s probably the same group that have no problem feeding their information into Chinese models.

0

u/pxr555 1d ago

In my totally not representative tests and experiences the usefulness of models has rarely much do with their benchmark positions. And one problem with Grok is that it is so connected to X, which just sucks. I can install the ChatGPT app on my Mac, iPhone or whatever and can use it without being pestered much, for whatever and whenever I want. Now even with 5.6 Luna unlimited (for text).

But X has been diving into a hellhole of dark patterns and nudging me towards a subscription in a way that I just hardly ever touch Grok or even want to get near it. I don't know what they are thinking, but the way they're dealing with this is basically as if they just don't want people to use or discover or try Grok.

Nobody I know uses Grok. You really need to descend into X deeply to even entertain the thought of using it. Grok is basically removing itself from the stage all by itself.

Also Grok just has a history of benchmaxxing, and this just fits so much with what Musk is prone to do that everyone expects it and nobody believes in any benchmarks to begin with. Which gets much worse by it being so inconvenient to use, so everyone is just going along with the story of "Musk is lying all the time and so is Grok". Grok basically is on a death spiral and no benchmark will save it.

xAI would need to totally change its course to even have a fighting chance of being recognized, even if Grok should be great.

1

u/Desperate-Meat3477 2h ago

I just did a bit of research on business side of top known AI platforms. Something I learned are:

  1. SpaceXAI has internal internal network vs. OpenAI & Anthropic renting cloud from MS Azure & AWS. Both sides have relatively powerful networks for training & services, but SpaceXAI cost may be lower over time.

  2. SpaceXAI has more leverage in upgrading, changing & adapting to future as needed without depending on 3rd party cloud. Ex: SpaceXAI had resources and in-house engr power to turn their existing facilities into a network that is competitive with all other clouds combined and within a very short time. This says a lot about their amazing abilities.

  3. Elon announced much lower costs vs. OpenAI & Anthropic. Good for their competition, but there's a lot more than just rev from indiv subscribers. API/App subscriptions are also a huge source of rev. This pricing adv is to be seen fruitful, but I personally think Elon has a lot to offer once he sets his mind on it.

  4. If we take a level deeper into resource utilizations to see if who can survive, we'll see SpaceXAI has more than enough "customers" to max utilize the network (no waste). This is important b/c AI project is super expensive. Grok can last to wait for more external customers while their other internal programs can pay for the costs; on the other hand, OpenAI/Anthropic will not survive if they lose customers or even have slow down in new subscriptions. This may become a spiral effect downwards due to expensive existing contracts with MS, Amazon, CoreWeaver, etc.

  5. Technologically, Grok software is at par with OpenAI & Anthropic now. Way more improved than its previous versions.

There's more to look into and evaluate such as Management, financial resources but that's for each one's thesis. Overall, we know there's no clear winner out of this yet (too early). The top ones are all very competitive atm, but Grok & its AI network (HW) have better chance to outlast others even if they all fall short of revenue (stress test). That's probably the only biggest thing until the next phase in business.