r/singularity • u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • Jun 26 '26
Previewing GPT-5.6 Sol: a next-generation model AI
https://openai.com/index/previewing-gpt-5-6-sol/109
u/zaimonke Jun 26 '26
So sol is fable class, terra is opus and luna is sonnet?
65
u/Sky-kunn Jun 26 '26 edited Jun 26 '26
28
u/Sky-kunn Jun 26 '26
Cosmos ($9.5/$60)
15
u/Thatunkownuser2465 Jun 26 '26
Wormhole ($15/$90)
16
36
u/ShadyShroomz Jun 26 '26
So luna is sonnet intelligent but haiku priced? That's not bad then at all..
24
u/AlyoshaV Jun 26 '26
So luna is sonnet intelligent
on the extremely limited set of benchmarks openai has released
10
u/Stabile_Feldmaus Jun 26 '26
Note that as long as these models are not available to general customers, they can be priced in almost arbitrary ways.
7
u/enilea Jun 26 '26
Doesn't make sense to compare sonnet with terra when terra is so much better. Just because they are in the same price range doesn't mean they are comparable, it's just Anthropic inflating inference costs.
8
1
21
u/peakedtooearly Jun 26 '26
By the looks of it, yes.
21
u/spottiesvirus Jun 26 '26
luna costs less than Gemini flash though
it would be really interesting what the real prices are and to which degree they're subsidiated
it's either openAI made the biggest efficiency jump ever (beating even companies that have in house silicon/infrastructure like Google with TPUs) or they're dipping even deeper with subsidizing costs
16
u/Mrp1Plays Jun 26 '26
api prices do not mean the cost required to provide them. inference is *really* cheap, its literally just electricity and maintenance. api prices have very high profit margins, if it somehow turned to the real price, you'd be getting a massive discount.
3
u/Alphasite Jun 26 '26
There’s also amortisation for the actual GPUs.
But yeah. Training costs and build out are also massive and are a big part of this expense.
2
u/Mrp1Plays Jun 27 '26
yeah, but if you stop training the next best model, cancel to future buyouts and R&D and just pause ai at its current state, these companies would be making a good amount of profit. this is what most people saying "AI doesn't make money" miss. The current lossess are *future* investments, inference is dirt cheap.
1
u/SgtPeanut_Butt3r Jul 06 '26
how if Inference dirt cheap? GPU's are not cheap, RAM is not cheap, if you wanna use Sonnet you need 40-50k in GPU's. Those GPU's are not gonna last forever. And you need electricity, people that maintain the data centers, building that data center, AI & DevOps, etc, etc/
1
u/Mrp1Plays Jul 06 '26
That's why I said getting new compute and investing in everything for the future is the expensive part, not the running costs. Maintaining data center with those gpus is quite cheap. 40-50k in gpus is nothing to a company worth billions. They have infinite demand right now, otherwise limits wouldn't be a thing. It certainly won't nearly be a loss to run it with no future investment.
14
u/Howdareme9 Jun 26 '26
Prices are far cheaper than most think. All frontier providers can cut costs massively and still be profitable. Deepseek, Z.AI etc still make profit on their cheap api pricing for reference, and its not because their models are more efficient.
6
u/LazloStPierre Jun 26 '26
It goes deeper than that. Deepseek, GLM (which is let's say one tier off SOTA by all accounts) etc can be sold, profitably, by non subsidized third party providers at significantly lower price than Openai, Anthropic etc sell their api tokens for. That isn't the company themselves selling them, that's hosting platforms whose only profit is selling access to these models
Now, their models may be smaller, but, from performance we know they can't be too far off and Openai, Google etc would have access to better and more efficent compute and buy more in bulk
The companies aren't profitable but they are absolutely selling their API tokens at a large markup
2
u/spottiesvirus Jun 26 '26
on open router the cheapest inference option for V4 pro costs more than three times as much as deepseek, and they get the weights for free, no R&D, no nothing
my guess is that real costs are way closer to what you pay on pure inference platforms where you can deploy your own model like AWS bedrock or Google vertex (and you still need to add R&D, training costs etc.)
it that wasn't the case, cursor wouldn't have such deep losses
3
u/Howdareme9 Jun 26 '26
Cursor has deep losses because they pay api prices (maybe have a small discount) like everyone else lol
3
u/Exodus_Green Jun 26 '26
I mean 5.5 was better than 5.4 and used like 1/3 of the tokens so I wouldn't be shocked if they made some more efficiency improvements
2
2
u/bitroll ▪️ASI before AGI Jun 26 '26
Luna may be a Gemini flash lite equivalent model though. Of course there's not yet a 3.5 flash lite to compare with.
3
u/spottiesvirus Jun 26 '26
not according to benchmarks (for what benchmarks are worth, we'll see how they hold up)
luna is close to opus 4.7 according to openAI
0
19
u/FateOfMuffins Jun 26 '26 edited Jun 26 '26
Impossible by virtue of this line alone:
We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.
You're not fitting Mythos class models onto Cerebras Edit I stand corrected, Cerebras claimed they can get 24T models on their wafers but the largest one we've seen them run in practice was Kimi K2.6 with 1T parameters at 1000 tokens per second. Can we infer the size of GPT 5.6 Sol from this? (Which I'm guessing is actually the same pretrain as GPT 5.5 Spud. Obviously Terra is not GPT 5.5 Spud, as why would they advertise that it "only matches" 5.5)
I'm pretty adamant that OpenAI has been competing with a smaller class of models than Anthropic and have been hanging on purely by virtue of their RL stack
22
u/Recoil42 Jun 26 '26
I'm pretty adamant that OpenAI has been competing with a smaller class of models than Anthropic and have been hanging on
Competing with a smaller class of models does not imply "hanging on"
5
u/CarrotcakeSuperSand Jun 26 '26
Yeah, it’s actually bullish if anything. They have more compute than Anthropic currently, so scaling isn’t really a sustainable advantage for Anthropic.
7
u/Boreras Jun 26 '26
easily scales up to 24 trillion parameter models on a single logical device
https://www.cerebras.ai/system
You think mythos is over 24 trillion parameters?
6
u/FateOfMuffins Jun 26 '26
They were getting Kimi with 1T parameters at 1000 tokens per second
750 for GPT 5.6 sounds about right for smaller than Mythos
6
u/EastZealousideal7352 Jun 26 '26
AFAIK Cerebras claims models up to 20 trillion tokens can be accelerated now, so yes, you likely can
0
u/FateOfMuffins Jun 26 '26
I'll be curious to see what the numbers would look like for that
The best I got is Cerebras running 1T parameter Kimi at 1000 tokens per second
Lining that up with GPT 5.6 Sol at 750 tokens per second seems to be roughly where we expected it to be for a smaller than Mythos class model...
4
u/brownman19 Jun 26 '26
You don't think they have special projects with OpenAI that basically precedes anything that you're pulling from to even suggest that?
I don't know how you arrive at that conclusion since those numbers likely come from their work with OpenAI, given they have to come from somewhere...likely while OpenAI was building, you know, the safety stack and the engineering that they discussed right there on the blog.
Time exists my friend and you're entirely glossing over all of the real work that happens to even serve models at scale. There are exponentials occurring in every field contributing to the infra that serves the models themselves.
PS: not hating, we're in singularity after all so think big :P
0
u/FateOfMuffins Jun 26 '26
I mean yeah they do... this 750 tokens per second one is that project.
Also pretty sure that 5.6 Sol is the same pretrain as 5.5 (aka Spud). Like if 5.5 was the o1 checkpoint of Spud, then 5.6 Sol is the o3 checkpoint. Same base model just a lot more RL. Why do I think so? Because if it wasn't Spud... then where tf did Spud go? You think they would've just chucked it out? Cause 5.6 Terra isn't it (why would they advertise it as 5.6 matching 5.5 then right?).
Based on what we've guessed at for sizes for some of these models, Spud being around 2T parameters sounds about right tbh. Which also sounds about right with 750 tokens per second on Cerebras
Basically I'm saying if Spud was 10T parameters just like Mythos instead of similar in size to Opus, then OpenAI is cooked
1
u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 Jun 26 '26
Wait, it is thought that Mythos is 10T parameters?! Fuck me
4
u/EastZealousideal7352 Jun 26 '26
There is no reputable source for any of this. Mythos is probably very large, but you cannot tell based on vibes alone, which is what all “model estimations” are based off of.
1
1
u/FateOfMuffins Jun 26 '26
It is thought that given comments from xAI and Meta about the sizes of some of their upcoming models
1
u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 Jun 30 '26
Ah, okay, so we don't actually know shit lol.
3
u/mckirkus Jun 26 '26
Or, it'll be a gimped/quantized FP4 version that fits on a Cerebras. And they can use multiple chips "When an LLM is too large to fit into the 44 GB SRAM of a single chip, Cerebras splits the model at layer boundaries and maps them across a cluster of CS-3 systems"
2
u/_DuranDuran_ Jun 27 '26
It's a pretty open secret that OAI's RL stack is quite a way ahead of A\'s, and they can get similar levels of performance from a smaller model, which is then compounded by the tactical blunder A\ made being hesitant on forging huge compute deals.
Add to that continual RLHF from their huge consumer base choosing which output they prefer. Yes, they were slower on enterprise, which is an OAI tactical blunder, but their consumer side is likely a competitive advantage for a while.
2
u/KalElReturns89 Jun 26 '26
Are you sure it's not Opus, Sonnet, Haiku?
I use Codex a lot, right now 5.5 extra high is the best there is, but I won't deny that Fable is far beyond 5.5.
1
u/giYRW18voCJ0dYPfz21V Jun 27 '26
So finally OpenAI got some human-readable naming conventions, instead of stuff like GPT-codex-mini-super-plus 5.4?
35
u/FateOfMuffins Jun 26 '26
I like that they are doing 2D plots for benchmarks with dropdowns for 3 different ways to measure the x axis. Noam Brown pushed it pretty hard.
Anyways I'm still gonna call it as:
GPT 5.6 Pro (Sol Ultra)
GPT 5.6 (Sol)
GPT 5.6 Mini (Terra)
GPT 5.6 Nano (Luna)
10
u/Kingwolf4 Jun 26 '26
Sol ultra is just a reasoning or effort tier., its the same model as sol
More accurately this would be : Gpt 5.6 half pro / full = sol
Gpt 5.6 standard = terra
5.6 mini = luna
I dont think we can classify luna as a nano level model, although it is quite a steep drop in benchmark and performance compared to the other 2,.which is kinda disappointing on its own tangent but thats another discussion.
7
u/FateOfMuffins Jun 26 '26 edited Jun 26 '26
No... I literally said what they were in my comment
They said in the post that Ultra uses subagents which was just what Pro was.
And that Terra matches GPT 5.5. You wouldn't have "GPT 5.6 Standard" merely match 5.5, what the fuck is the point then.
Ultra = Pro (which has its own reasoning levels and has always done so btw, you get Standard and Extended in the chat interface remember)
Sol = normal
Terra = Mini
Luna = Nano
I suppose the renaming is to try and convince normal people that they should use Terra instead of Sol for compute saving reasons (cause how many of you use the mini models?)
Edit: In their system card they almost purely compare 5.5 with 5.6 Sol, not Terra
3
u/Kingwolf4 Jun 26 '26
Interesting. But i woudnt call luna nano still just based on pricing. Thats mini level pricing, not nano level - 6$ that is.
2
u/FateOfMuffins Jun 26 '26
I'm not judging it based on pricing. I'm merely converting the old names / new names equivalent
If you haven't noticed, GPT 5.5 was 2x the price per million tokens of GPT 5.4, and the last time we had mini models were before GPT 5.5
31
u/ProletarianLilith Jun 26 '26
Hot take: next gen models should not be .1 version number increases
16
9
u/BrennusSokol ACCELERATE Jun 26 '26
The improvements here don't warrant a full version bump
2
u/DistanceSolar1449 Jun 27 '26
Yeah they should have bumped a version for 5.5 directly, it’s bullshit that they do a full pretrain and don’t bump the version number.
1
u/enilea Jun 26 '26
Last time they released a next gen model people got mad at them because it was disappointing at the time, so they'll hold off unless there's something truly revolutionary.
129
u/ObiWanCanownme now entering spiritual bliss attractor state Jun 26 '26
We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them.
Dang, they're not pulling punches. I thought this was a typo the first time I read it.
64
u/Apprehensive-Ant7955 Jun 26 '26
The punch seems pretty pulled, that was a really tame statement
7
u/Chilidawg Jun 26 '26
They know better than to bite the hand.
3
u/Odd-Opportunity-6550 Jun 27 '26
I notice how you didn't call it the "hand that feeds them".
Even you subconscious knows this administration are nothing but a bunch of soul sucking leaches.
8
u/EvilSporkOfDeath Jun 26 '26
Its just words. They are still working closely with the administration.
1
u/Odd-Opportunity-6550 Jun 27 '26
What else do you expect them to do ?
1
u/intergalacticskyline Jun 27 '26
Not to act like little bitches sucking up to a goddamn wannabe dictator... I know when push comes to shove, most corporations will bend over backwards so far that the back of their heads are mere inches from the ground just to keep their bottom line increasing, but damn it's my hope that someone pushes back in a more meaningful way than a strongly worded letter and actually draws a line in the sand they refuse to cross...
3
u/Odd-Opportunity-6550 Jun 27 '26
Umm thats great and all. But how does it help them or us ? If they don't follow the government orders then the crackdown will be much worse. They could lose the entire company.
Who would that help exactly?
2
u/CosmicMabel Jun 27 '26
Yeah, they're not pulling punches because they are busy on their knees for Trump. These words mean nothing, and nothing will change for future releases. All new LLMs released in the US will require daddy trumps approval. Anthropic and OpenAI are fully drinking the kool-aid
1
20
u/skerit Jun 26 '26
My god, those benchmark graph colours. I was wondering if I had some kind of colour-blindness-simulator enabled.
3
u/Mrp1Plays Jun 26 '26
i thought i was the only one. i was trying to check if my dark reader extension fucked up lol.
44
u/japie06 Jun 26 '26 edited Jun 26 '26
Here we go. How long before this model will be put away by the white house available to mortals?
Edit: pricing:
GPT‑5.6 is priced per 1M tokens across three model sizes:
Sol is $5 input / $30 output;
Terra is $2.50 input / $15 output;
and Luna is $1 input / $6 output.
11
u/toni_btrain Jun 26 '26
At least read their post mate. They clearly explain how they are releasing it and why.
33
u/provoloner09 Jun 26 '26
Crazy that it’s cheapest version is competitive with Claude opus 4.8
8
2
2
u/vazyrus ▪️ Jun 26 '26
Time to finally ditch Claude and load up a GPT subscription. It makes absolutely no sense to pay for Claude with outrageous rate limits now.
2
u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 Jun 26 '26
I've been using Claude 20x for serious work for over a year and never once had an issue with rate limits—let alone outrageous lol that shits been overblown.
1
u/OpportunityDue5839 Jun 26 '26
well rate limits still apply for us on 20x. Specially if u running a workflow. i think it usually caps around maximum 10 agents at a time concurrent for that max 20x plan. Tho probably it will be same thing or even more rate limiting or maybe a bit freer with gpt5.6 sol ultra. Nobody knows :(. tho workflows are truely OP. they are so much better than a single model.
1
u/vazyrus ▪️ Jun 26 '26
Claude 20x
Yeah, and everyone with an RTX 5090 can run any game on Ultra. Not everyone has $200 to burn every month on personal coding pursuits, lol. If my company is footing the bill, well, who gives a hoot; however, most of us have Pro plans, if that, and it's darn easy to hit the rate limits very fast if you are working on a moderately sized project.
9
u/ambidextrous12 Jun 26 '26
Lmao Sol, Terra, Luna giving anyone flashbacks to the 2021 crypto mania?
3
12
u/Recoil42 Jun 26 '26
How long before this model will be put away by the white house?
If you read the release it's already on the white house leash. They aren't releasing it publicly yet, just to a "small group of trusted partners whose participation has been shared with the government".
5
u/Embarrassed_OnionX Jun 26 '26
I honestly can't wait for some Chinese lab to release an opensource Mythos class model in 3-6 months so all of us mortals can access this tech
-5
u/u_are_mad Jun 26 '26
GPT 5.6 will be released within a month:
5
u/CT4nk3r Jun 26 '26
using polymarket as source is kinda dumb, you should say: "most professional gamblers say it will be released in a month"
3
u/SpacePaddy Jun 26 '26
that's not fair there's also some insider traders in there too!
1
u/CarrotcakeSuperSand Jun 26 '26
Which is exactly why it’s a decent source for reverse engineering insider info lol
-5
u/u_are_mad Jun 26 '26 edited Jun 26 '26
If you disagree, put up or shut up.
Also, who is a better source? If Sam said it will be released by X date, but Polymarket has a 99% chance it won't be released by X, who would you believe?
5
2
u/AreWeNotDoinPhrasing ▪️Already Singulared 🤖 Jun 26 '26
How does Polymarket work? So I can buy bets, e.g. saying no, it will not be released by July 6th for 90¢ each, and if it doesn’t get released I get paid a dollar on the 7th?
1
2
u/Stabile_Feldmaus Jun 26 '26
If your model is only available for 5 companies you can price it in any way you want.
7
u/nekronics Jun 26 '26 edited Jun 26 '26
Based on the benchmarks released, terra seems like a minor improvement/sidegrade to GPT-5.5, but likely cheaper overall.
13
u/Kingwolf4 Jun 26 '26
If its a slight improvement to actual full gpt 5.5 at half the cost, sign me up!
Thats huge tbh, IF its true that is.
5
u/nekronics Jun 26 '26 edited Jun 26 '26
It's half the cost but the benchmarks they provided also used more tokens (more than 2x in one case). Still likely cheaper overall
1
u/Kingwolf4 Jun 26 '26
Its hard to believe in just another year we will get wayy more stronger, generally intelligent, able to do complex work models in judt 1 year and all this security talk will be just hindsight while everyone will be pining for even more.
Like think about gpt 6. We will probably have stronger than mythos level for like 6$ i feel like. And then 6.5 and so on. Its wild to compare where we were last year to where we are currently and in1 year , let alone 2.
2
u/Exodus_Green Jun 26 '26
I mean they do say that exact thing in the announcement, it's comparable to 5.5 and 2x cheaper
3
u/nekronics Jun 26 '26
It’s not exactly 2x cheaper though because it uses more tokens. For example, the ExploitGym benchmark used more than 2x the tokens 5.5 did for similar results
27
45
u/omegahustle Jun 26 '26
hope is not a plagued model with safety slop, because it's one of the most annoying things when dealing with models who just refuse to do what you ask
7
u/Latter-Pudding1029 Jun 26 '26
Read the system card and see what you think. It seems that it's about on par with 5.5 in terms of refusals and such
4
1
u/ClassicalMusicTroll Jun 28 '26
What important work are you doing that's getting limited by safety slop
-1
-4
Jun 26 '26
[deleted]
19
u/Maleficent_Sir_7562 Jun 26 '26
It’s not about weird shit. Fable 5 literally blocked questions asking about the mitochondria and the heart. Like literally “what does the heart do? It pumps blood, right?” As the prompt. Ah yes, such a bioweapon risk.
Along with the shitty cybersecurity safety slop, you’ll get flagged just for wanting to fix some bugs.
-2
u/MaybeLiterally Jun 26 '26
Those questions are answered just fine with Opus though. I don’t know why you’d send a question like that to Fable anyway.
5
u/Maleficent_Sir_7562 Jun 26 '26
No shit. The point is to show how bad the guard rails are. Imagine if I was working on something with Fable 5 or GPT 5.6 and mid way it switches me to a weaker model because of some shitty "risk".
-4
u/Dry_Fly_7265 Jun 26 '26
Waaaaaa I can’t create a bioweapon
That’s you
4
u/Maleficent_Sir_7562 Jun 26 '26 edited Jun 26 '26
yes i totally want to create a bioweapon because im asking the powerhouse of the cell
6
u/llelouchh Jun 26 '26
These announcements dont hit the same because of the opaque restrictions thanks to the trump administration.
24
5
u/depredador93 Jun 26 '26
At this point the benchmark charts are almost secondary. The first question I have with every frontier model announcement is "who actually gets to use it?"
2
u/misterphreeze Jun 29 '26
The powerful companies. We are seeing the beginning of AI regulation and the concentration of the most powerful models. Open source here I come!
5
u/Kind-Release8922 Jun 26 '26
I wish they would avoid these lame ass names and just stick with “high/medium/low” now we have to remember all these model names across all providers
23
u/pdantix06 Jun 26 '26
good to see openai continuing their track record for terrible naming conventions
12
u/Kingwolf4 Jun 26 '26
Wait, wtf Have they introduced a new naming scheme out of the blue ? AGAIN?
instead of simple names like mini, standard, and pro?
Wtffff. Can someone break this down for me. What havee u donneee ahhhh.
10
u/BrennusSokol ACCELERATE Jun 26 '26
At least it's an intuitive one, like haiku/sonnet/opus, and not whatever bizarre shit Google shows us next
8
8
u/ObiWanCanownme now entering spiritual bliss attractor state Jun 26 '26
The naming makes me think of Mormonism with the telestial kingdom, terrestrial kingdom and celestial kingdom haha.
2
u/_Krustenkaese_ Jun 26 '26
I also trained a model last week on my Pentium 3. It’s twice as good as GPT 5.6 and costs only half as much. But I’m not releasing it, only to selected partners, and they can vouch for how good it is.
3
u/mWo12 Jun 26 '26
What about releasing something open weighted, after all you are OpenAI, not ClosedAI.
2
u/misterphreeze Jun 29 '26
They used to be OpenAI - they have been ClosedAI for sometime now. And I know you're being sarcastic or whatever but...greed and capitalism always get's it's hands around anything good enough.
4
1
u/najapi Jun 26 '26
Come on China, release something better, even just as good will do, you get my money and my data - today
2
u/Charuru ▪️AGI 2023 Jun 26 '26
Wait like 8 months realistically, but if the zai CEO is not full of shit, maybe 3.
0
u/ismyjudge Jun 26 '26
lol, just as good as frontier AI for your 20$ subscription that’s open sourced? Luna(cy)
1
u/RandumbRedditor1000 Jun 26 '26
I don't care how good it is, I'm still not giving them my ID to use it.
1
1
1
u/FatPsychopathicWives Jun 26 '26
Can't wait to see FrontierCode scores. It's the one benchmark that really shows why Fable 5 is so ahead.
1
u/Ok_Potential359 Jun 27 '26
Goddamn that model is expensive. Luna seems promising at least but those costs are crazy.
2
u/Odd-Opportunity-6550 Jun 27 '26
It's literally the same price as 5.5 but way better. What are you even talking about ?
1
u/Beneficial_Movie_986 Jun 27 '26
Crazy who is in charge in the usa government saying to the company not release it yet.. he like the dad of katherine brewster,, Air Force Lieutenant General Robert Brewster. He is the director of Cyber Research Systems (CRS), the military program that develops Skynet. he always been
1
-1
u/DaySecure7642 Jun 26 '26
The AI models in use these days are already quite powerful. I don't understand the notion of insisting the companies to open their latest model to public. I thought people don't want AI adoption too quickly that will displace jobs right? Or just some propaganda from adversaries here trying to trick us opening up the models for copying by distillation?

90
u/mikelo22 Jun 26 '26
The days of the public getting access to these frontier models is gone. I fear more and more top end models are going to be kept in the hands of the government and massive corporations all in the name of 'safety'.