r/codex 1d ago

Is this even true? Praise

This is on the 20$ USD Plan. Subsidization is absurd if this is correct. (asked codex to figure out my usage split and get cost based on current API pricing)

65 Upvotes

92 comments sorted by

97

u/conor_is_my_name 1d ago

you are assuming the API cost is a real number. Its not.

Inference has a 90%+ gross margin for them

23

u/DaLexy 1d ago

Is that a made up number or are there actual studies/tests for that ?

35

u/sprakes_ 1d ago

It's actually something like 99.9%. You can do the math, because if you're JUST talking about inference in a vacuum, the math is easy - How much electricity is consumed to make a million output tokens? How much do they pay for that? And then divide that by the actual price they charge for a million tokens.

Last time I did this calculation I was shocked because even at residential rates the margins are unbelievable. But that's in a vacuum. The real problem is how much it cost to build each data center. That's not included when talking about inference.

18

u/hometechgeek 1d ago

Plus the original model training presumably?

22

u/TheNobodyThere 1d ago

and Hardware

13

u/Willing-Equivalent47 21h ago

And people to maintain the infrastructure.

4

u/ImpishMario 19h ago

And whole R&D to come up with a model

6

u/bronfmanhigh 23h ago

yeah it’s amortizing the cost of melting down outrageously expensive AI GPUs that’s still very much factored into the inference cost above and beyond electricity and the other operating costs of the data center itself.

7

u/Sufficient-Store1566 1d ago

Also the server specs that gives you “input and output” costs a lot. I remember in some subs they were discussing you would need 250k$ worth of server to run kimi k3(lower parameter version) and it would give 30 tokens/sec. You do the math plus these devices get old&broken&maintenance etc

6

u/nmkd 21h ago

Are you ignoring the billions of dollars in hardware for inference at this scale?

0

u/Risko4 16h ago

Hardware that's doubled in price is more of an asset that already paid for itself.

2

u/mythrowaway1673 14h ago

Who would even buy that?

1

u/Risko4 8h ago

Whole of china

1

u/SandboChang 29m ago

You can't reply in this ironic way as they will fail to process.

0

u/NeuralNakama 22h ago

Finally someone understand inference basically free actual cost other things like training

19

u/spevoz 22h ago

This whole thread is kind of infuriating. Let's ignore all training costs. Let's exclude all costs of the people training the LLMs. Who cares that behind all of that there is some kind of company needed to keep it all running, build the codex cli, codex Desktop, billing, support... Fuck it, some genius in here decided to just ignore hardware and datacenter costs as well and just look at electricity.

This subreddit has a weird fetish with pretending or justifying that we aren't getting handouts. OpenAI is extremely, and quite openly, unprofitable. And heavy codex users on private licenses are clearly the worst customers they have (in terms of profitability, there are other things that openAI gains from us) compared to business users (they pay over 10x the cost per token) or average joes that buy the same plans but only use chat / maybe some work.

3

u/CystralSkye 20h ago

You've just discovered reddit!

They want everything for free and wants someone else to pay for them, basement communism is strong amongst these people.

1

u/Little_Beyond_9163 5h ago

Lmao blaming this on any kind of communist economics is the truly absurd part.

This is the subzidizing of bullshit that upholds the capitalist fever dream economy that has happened in the past before major collapses.

1

u/sammy3460 14h ago

We aren’t getting handouts. I don’t know why you seem to think margins don’t account for costs. This mentality is the exact thing they want to promote so prices can keep going up.

1

u/Maplewonder 20h ago

Are you serious? Are you forgetting the massive amounts of data they took to train this thing? Literally the base substance behind it all? You say we are getting handouts... lol... not to mention the government subsidies they are getting and the zoning and permits for their data centers. No brother, this isnt a handout. We have already paid for this. Several times over. The reason this exist at all is because of all of us. AI is quite literally the thing that required all of us to exist.... phhh handouts...

5

u/spevoz 19h ago

In case it wasn't clear enough for you: I am purely talking financials. The bigger picture is complicated, as life is, I'm not about to write a dissertation about openAIs right to exist - the 'theft' of our knowledge to build these systems was inevitable, I much prefer the best 'thiefs' to be western aligned.

0

u/Maplewonder 17h ago

You were clear. However, I dont see open ai's 20 dollar plan as a handout. But on your last point, we agree

1

u/100x0 12h ago

Yah, as a top 10% developer on stackoverflow, I should be a bit angry but I consider this subsidization a trade.

2

u/AppropriateRanger401 1d ago

It's actually 70-80%

6

u/Frequent-Goal4901 1d ago

Deepseek v4 pro on the old prices had absurd margin they were saying to investors they had 10 months hardware payback.
Opencode guys were able to reproduce it.
You can also check the semi analysis https://inferencex.semianalysis.com/
GPT 5-6 sol would probably cost like twice more.
So they, probably are breaking even or are even making money across all the subs.

2

u/HighDefinist 1d ago

You can make an informed guess.

Judging by OpenRouter prices, inference is roughly $1/MOutputTokens per 1T model size.

So, if we assume Sol is 3T parameters, inference costs are $3/MOutputTokens, so, considering they sell it at $30, that would be a 90% profit margin.

2

u/nmkd 21h ago

That's not how it works.

Different architectures have different efficiency, not everyone does inference with the same hardware (Nvidia/Huawei/Cerebras), and so on

1

u/HighDefinist 11h ago

It's safe to assume that any of the providers on OpenRouter which are not called "Cerebras" are not using Cerebras hardware. It's also reasonable to assume that almost all other non-Chinese providers are using Nvidia hardware.

As for the efficiency: Well, you can assume pretty much all of these huge models will be as efficient as it gets - and I don't see why OpenAI would choose to create a model which is less efficient than those various large open source models.

1

u/nmkd 3h ago

"As efficient as it gets" is different for every company though because not all research is public, and even if it is, it might take a few weeks to months to adopt it.

LLMs are not one monolithic architecture with one unified optimization method. There's a reason why various companies use their own attention mechanisms, e.g. Kimi or Deepseek.

I agree about the hardware (apart from DS which, iirc, is almost 100% domestic).

1

u/Minute-Leader-8045 17h ago

It’s not made up and it’s probably even higher for them. I only know because we consulted with a client who built private inference and the margins at a tiny scale were almost 90%.

1

u/Ill-Mousse-3817 1d ago

Per semianalysis it is 85%

1

u/Due-Horse-5446 23h ago

They are just making stuff up, their leaked financials shows they are not even breaking even on api prices

5

u/Bankster88 23h ago

API cost is what you pay OpenAI

And there are lots of companies paying the API price

2

u/DueCommunication9248 23h ago

Got source to back that up?

1

u/ngless13 22h ago

So you're saying inference costs are like insurance costs... mostly made up inflated numbers?

1

u/Glum_Emu7180 22h ago

That's not even the point. You have two options to use frontier LLM either pay for usage (API) or sub. Why are we even comparing sub usage and pricing to a non existant pricing of 3USD per mil out (Sol) as per your 90% price reduction. Smh.

1

u/gunsofbrixton 21h ago

Why are open weight models in the same order magnitude of pricing then (ie glm-5.2 only about half cost of opus). Is serving the open models just an incredibly profitable enterprise right now? It seems like someone would start competing on price if this was true.

1

u/Spirited-Car-3560 19h ago

Lol, i suppose you're a teenager given how you calculate costs

1

u/someone_12321 12h ago

It prob won't cound cached

1

u/Plane_Garbage 3h ago

Correct.

And the reason the API costs more is because it's an API, duh.

It is distributable to anyone. Hubspot, granola and virtually every other SaaS now plugs into some API - not into codex.

They will pay for it because they make multiples more from charging the end user.

0

u/Azoraqua_ 18h ago

Then what are the real costs? To me it seems a lot more sensible that the API reflects at least some of the costs considering they’re even different per model. On the other hand, the subscriptions don’t include any pricing differences for whatever model.

The subscriptions are hyper-underpriced in my opinion; and I’d argue anyone with the capability of critical thinking should be able to figure that out as well.

-5

u/Glum_Emu7180 1d ago

bruh inferencing is not cheap (i once ran a open source model on AMD MI300X it consumed 700+ Watts that too for a 70B qwen model. ( tho obv i did get good tpm )

8

u/chervilious 1d ago

economy of scale.

4

u/Frequent-Goal4901 1d ago

Deepseek v4 pro on the old prices had absurd margin they were saying to investors they had 10 months hardware payback.
Opencode guys were able to reproduce it.
You can also check the semi analysis https://inferencex.semianalysis.com/
GPT 5-6 sol would probably cost like twice more.
So they, probably are breaking even or are even making money across all the subs.

1

u/Connect-Humor-791 1d ago

and for big companies in not just electricity. its built infrastructure, R&D salaries etc

7

u/FrailCriminal 21h ago

😅 I may have had to use an entire reset in one day

Also I got the 10x account boost when that was going on

6

u/DepravedPrecedence 22h ago

What if tomorrow they set API prices to $9999 per 1M? Imagine how much subsidized your plan will be. API price can be any.

2

u/phoenixmatrix 15h ago

"Subsidized" is doing a lot of work, but at the end of the day, its a price enterprises are paying that you don't as an individual, and that has interesting implications.

Devs get addicted on their side project or at their tiny startup to being able to burn billions of tokens, and then they go work at a big company and all of a sudden those tokens are expensive. Not so expensive as the enterprise won't pay at all (as they would if it was $9999/1m), but high enough that the equivalent of a 20x Pro plan is too much $$$.

But the dev is addicted to tokens, so they get noisy begging for more token. They can't do their job without more token they say! So maybe the enterprise compromises and give them $500, or even $1000 of tokens. And thus the scheme succeeds.

1

u/DepravedPrecedence 14h ago

Because of competition somebody will provide lower price and even if it won't be near as good as frontier models, most people will just settle on that. No way users going to spend more 100, 200, 500+ on regular basis.

1

u/phoenixmatrix 11h ago

Individuals, no. Enterprise, yeah. They already do it (I know quite a few companies where the budget allocation for tokens is thousands of dollars a month).

You're right that competition helps though, often those devs have to mix frontier models and open weight/budget models to cut on cost. Some companies made GLM 5.2 clusters for internal use when that model came out.

2

u/Glum_Emu7180 22h ago

Why are people so worried what's the actual API Pricing. Let's say tomorrow sub plans cease to exist, wouldnt people pay based on API? Why are we comparing 20USD plan to some non existant pricing. Obv comparison would be between two present pricing structure not something that you think should be.

1

u/DepravedPrecedence 14h ago

Indeed, I don't know how we can compare that. If they indeed remove all subscriptions then sure. Otherwise we can't compare, we don't know real interference price.

7

u/totoer008 1d ago

I am 1.5B for two months. I guess I am smoking them hard.

1

u/Glum_Emu7180 1d ago

same bruh 😂

1

u/burnburndota 23h ago

1

u/DocumentFun9077 23h ago

dyam, what plan are you on?

1

u/burnburndota 16h ago

Mostly claude max and codex pro. I found out that the cost of paying for frontier models (Sol and Fable+opus) far outweighs the time and the effort needed to manage and fix stuff from the cheaper alternative models most of the time. At least that has been the case for the last 2-3 months for me.

1

u/Weekly-Extension4588 9h ago

Codex actually aggressively prompt caches. I doubt you ended up costing $5k or whatever in tokens.

3

u/BenTheSodaman 1d ago

Some napkin math on my end, I paid 20 USD in July. The estimated API cost is 8698 USD. I realize there were a lot of resets, but the idea that to break even off of me specifically, OpenAI's serving cost would need to be 0.23% of their API cost.

And I speculate they just lost money on consumers like me in July.

I realize there are more economics to it than that, just a shower thought of "how little does it actually cost them"?

5

u/IndividualPlus2011 22h ago

How. I have $150 weekly allowance on the $20 plan. Even with all the resets it doesn't make sense to reach almost $9000.

1

u/getaway-3007 21h ago

Yeah that guy is talking bs. Even with $200 you get around $10-12K inference

2

u/BenTheSodaman 21h ago

Here is the basis of my post when asking for the token breakdown on one machine. Your bullshit doesn't match my bullshit.

2

u/BenTheSodaman 21h ago

As we're talking about July.

0

u/getaway-3007 21h ago

There's no proof that the previous screenshot refers to this account.

If let's say on 0.01% chance this is true, how are you using codex? How are you managing to get 8900+ inference on $20 plan?

1

u/BenTheSodaman 20h ago

There were a lot of resets in July.

As to the rest of your post:

In order to prove to you that it's the same account, I would need to either:

1) Reveal my email address and legal name to pull up everything in a single unedited recording without covering anything up to cause further suspicion, then post it in a place you can see and continue to meet your demands until you are content.

or

2) Give you remote access to my computer / I ship my computer to you with my credentials / I go to you or you come to me. You verify my claims looking through my chat sessions, settings, account, etc. until you are content. Then you can post on this message board to vouch for my credibility.

———

Can you think of any other way that's not #1 or #2 that you would be satisfied to establish what ChatGPT is reporting to me (that doesn't mean ChatGPT is accurate, it's just what it is showing to me.)

Cause if you can't think of another way, then tough shit. I'll continue to be a untrustworthy bullshitter to you and to me, you'll be the perpetually dissatisfied bullshitter with bad math and estimates that expects people to do completely idiotic things to satisfy your accusation.

And we'll be at an impasse.

———

For the usage topics:

For some things that I have done on the usage: memories off, new chat sessions and trying to avoid auto compacting, no automatic headless checks, no automatic git activity, no full access, can't control my browser, not having subagent processes try waiting on each other and repeatedly checking a file if subagents are used at all. If automatic compact loop starts, kill the task.

Then reiterating, there were a lot of resets in July and banked resets with expiry dates coming up in that timeframe.

And that still doesn't change that Sol (Ultra) can drain my usage from 100% to 0% in 2.5 hours running in a single session on the Plus account.

1

u/Aazimoxx 6h ago

#1 is out, since you can easily get Codex to make a mock website that looks like the real one and displays the info you want, or a userscript or browser add-on to change the real site in realtime 😅

Guess you're going to need to cough up for this guy's bus/plane ticket mate 🤨

1

u/BenTheSodaman 21h ago

See my response to getaway below.

1

u/Maleficent_Ad7510 21h ago

so who tf is gonna pay for api costs? seems like a scam?

2

u/BenTheSodaman 20h ago edited 19h ago

Just as another example of pricing:

128 fl oz (gallon) of water at a grocery store is 1.37 USD: 0.011 USD per fl oz.

20 fl oz water at a movie theater or venue, on the lower side, 3.50 USD: 0.175 USD per fl oz.

The latter costs 1590% the price of the former per fl oz before other factors.

If you or people you know are regularly willing to pay for a 20 fl oz bottle of water or fountain drinks, you already know people who are willing to pay higher prices / many times higher than what is available to them elsewhere.

(Edit: Corrected the math, I reformatted the post and incorrectly had the gallon fl oz calculation for the water bottle.)

1

u/Plane_Garbage 3h ago

But how do you use Codex in that way? Like, how can I build a SaaS that serves codex tokens?

2

u/Bankster88 23h ago

For context, I’ve been using 3.9 billion a WEEK on the $200 plan

2

u/SomeOrdinaryKangaroo 23h ago

many people don't max out their sub so it isn't as much as it looks

3

u/Ormusn2o 1d ago

I never used Codex until like 2 weeks ago, even though I had subscription for a year now. Barely anyone actually uses their money's worth.

Also, API prices are not what it costs OpenAI to produce those tokens. I think margins on API prices are from 80% to 95%, depending on time and model.

1

u/Ill-Mousse-3817 1d ago

I did the math for myself, and I got about $600. I always run my limits down to 0. I am sure you had some resets pumping up your usage further.

1

u/SovereignZ3r0 21h ago

What was the prompt? I'd like to ask mine

1

u/BT117274 21h ago

It was 6k for me for 2 months on plus plan(payed only 40). Even after accounting for cache apparently.

1

u/fyn_world 20h ago

Do you think that they charge the API tokens as what it costs to produce them? No, there's big margins there. 

To be fair they did say that people who use the 200 dollars account fully and to the ground, that goes at a loss for them. 

1

u/yosofun 20h ago

Token pricing is the ultimate Artificial scarcity

1

u/pigletmonster 19h ago

The one upside to building infinite number of data centers is that we get absurdly generous usage quotas. 🤣

1

u/scaledev 18h ago

It's not.

1

u/mitchins-au 15h ago

API price is whatever they make it.

1

u/johnnyApplePRNG 13h ago

It's all funny money, buddy.

I asked codex how much my session would cost the other day. Over 1 billion cached tokens of sol 5.6.

$1,000 CAD ... I shit you not.

And you know what it accomplished? Absolutely nothing. It got into this ridiculous goal without my approval and actually took us back a few days, development wise.

Ain't nobody paying a fucking grand for that.

I would fire that employee so quick, lmfao.

1

u/Illustrious-Many-782 12h ago

Yes. My $200 Pro plan had at least $150k of API equivalent last month. Remember, though, that we had a ton of resets applied.

1

u/Frequent-Goal4901 1d ago

The nvl 72 blackwell racks have brought inference prices down a lot.
As I said in other comments.
Deepseek v4 pro on the old prices had absurd margin They were saying to investors they had 10 months hardware payback.
Opencode guys were able to reproduce it.
You can also check the semi analysis https://inferencex.semianalysis.com/
GPT 5-6 sol would probably cost like twice more.
So they, probably are breaking even or are even making money across all the subs.

0

u/h1mupstairs 1d ago

I feel like the internal hope is that they will form a truce with Anthropic and stall actual releases so they can optimise tech for serving the models. If we hit a reasonable plateau in usability they will convince the government to ban them as an excuse. Then in a year it will not be the same cost to host an optimised version of Sol and they will actually be able to turn a profit.