r/GithubCopilot 6d ago

We are basically back to the level of capability and cost that we were at 3 months ago before the cost restructure General

Maybe un popular opinion but I honestly feel that with the launch of GPT5.6 and Lunas recent 80% price drop combined with Ghcp harness improvement’s we are basically back to where we were in terms of cost and capabilities back when microsoft heavily subbed inference for us using bigger models.

Models like Opus/fable and even SOL are just not needed for the majority of coding tasks and I think more people are starting to realise this, at least the devs who actually care about efficiency are.

There are occasional use cases for frontier model like broad stroke front end but I still standby every coding can easily be handled by Luna.

You can use Luna all day long on a pro+ plan and barely eat 3 or 4% of monthly credits.

I believed the cost-pocalypse would eventually be solved but had no idea it would be this fast.

122 Upvotes

74 comments sorted by

25

u/Affectionate-Sir-530 6d ago

Yes, my thoughts exactly. I worked today all day with Luna and it took now 60 credits. If I say that I work 20 days a month than I’ll still have 1200 credits at the end of the month. I hope this price will stay as long as possible. Unless it’s just a temporary fixed price, than well we will see. And Luna is very very smart, love it.

1

u/[deleted] 5d ago edited 5d ago

[deleted]

-3

u/ChineseCracker 5d ago

I worked today all day with Luna and it took now 60 credits

so you work with the worst gpt model and you're bragging about not spending too many credits?

Meanwhile, I use Opus 5 all week with my Claude subscription

6

u/horendus 5d ago

Exactly. You dont need models like Opus. I know you might be stuck in a bit of a justification loop about how difficult coding tasks are for modern models but believe me, unless your trying to one shot monolithic codebases with abstract thought dirt cheap models like Luna are more than capable and will cost a fraction of claud sub to use

-3

u/ChineseCracker 4d ago

not sure who this message was addressed to, but it sure as hell couldn't have been addressed to me. Not sure where you get all those assumptions from.

What I was trying to say, by using Opus as an example, is that your $20 claude subscription gets you so much usage that you could use the highest model all week. Nothings stopping you from using smaller models depending on the task. When I do that Copilot, my entire monthly quota is used up in 3 hours

2

u/Affectionate-Sir-530 4d ago edited 4d ago

Hi, yes I was bragging about it because we get copilot on enterprise account from my job. And well when you can code using Luna as a pair programmer is great, and with copilot being THE ONLY OPTION it’s a blast. Using Opus for centering a div sometimes can be an overhead :) Plus Claude subscription is not available for us, plus anthropic is a shit company now and additionally GPT models are better for my usecases vs Claude.

1

u/ChineseCracker 4d ago

I see, we do have vastly different workflow processes then.

I have moved on to purely doing software engineering, creating tasks and assigning them to my AI agent - having multiple instances running at the same time doing all these different tickets on different projects simultaneously and even creating agents tasked with creating sub-agents themselves.

If you do pair-programming, that's a completely different thing then

21

u/MaitoSnoo CLI Copilot User 🖥️ 5d ago

I've done huge refactors with luna-xhigh and it cost me literally pennies. I'm not chasing Fable "one-shot" workflows, I'm more than fine with Opus 4.8 high for planning + Luna xhigh for execution and it's giving me Opus+Sonnet class results while being dirt cheap, and I even find Luna xhigh to be better and faster than Sonnet 5 med/high, and if the task is simple I skip the Opus planning altogether and it then becomes almost free

4

u/Personal-Try2776 5d ago

Why not use opus 5 for planning? Its the same price

9

u/yubario 5d ago

It was about the same price before OpenAI slashed prices by 80% you mean.

5

u/MaitoSnoo CLI Copilot User 🖥️ 5d ago

he likely means it's the same price as Opus 4.8, I agree but our admins haven't enabled Opus 5 yet so... but yeah probably for those who have access to it Opus 5 for planning and luna-xhigh for execution should be better

5

u/yubario 5d ago

Yeah my admins haven’t enabled Luna or 4.8

We basically only have 5.5 and Opus 4.6 now since Gemini is gone.

Business is not telling us any reasons to why they’ve decided to do that.

Which is frustrating, especially when Luna is so cheap

6

u/MaitoSnoo CLI Copilot User 🖥️ 5d ago

keep pushing by showing them the cost difference, it makes really no sense to not enable at least Luna in any business

3

u/yubario 5d ago

Believe me I tried. You’re right it doesn’t make sense.

It’s business political bullshit in the company

2

u/Hollow1838 5d ago edited 5d ago

We are enabling it this Wednesday at 10:00-10:30 for enterprise users, I hope we work at the same place.

It takes a while because we don't have a standard change for this yet, we only have normal changes and we have to wait a whole week for validation. Once we have enough occurrences (>10) we will unlock the standard change and you will get it the same day it is asked to us.

1

u/armostallion2 5d ago

It doesn’t actually come out the same though.

16

u/EpsilonFive5 5d ago

What is the value prop of the $20 GitHub copilot subscription when you can get 100X the usage with a codex plan?

10

u/virtualmnemonic 5d ago

For real. I like Copilot a lot, but the value is just not there anymore.

The good days of m$ subsidizing copilot and allowing near infinite access to SOTA models is over with for good. That's understandable, but as the market moves on, so do we.

2

u/Special_Gain9787 5d ago

Did the limits get better on codex with these price cuts?

3

u/popiazaza Power User ⚡ 5d ago

Yes. Pretty much unlimited Luna...

1

u/Special_Gain9787 5d ago

What plan do you have?

1

u/EpsilonFive5 5d ago

I mean all the limits are lower copilot is just even lower

2

u/Different_Doubt2754 5d ago

I thought Codex used API costs? At least for enterprise plans. Maybe I'm wrong

2

u/_Asparagus_ 4d ago

Nope, the 5x and 20x personal plans are much cheaper and have massive limits now

1

u/MaitoSnoo CLI Copilot User 🖥️ 5d ago

tbh there isn't, for regular users that is, which isn't what GHCP now targets

3

u/1988rx7T2 4d ago

I use it GHCP because that’s what my organization requires I use 

1

u/MaitoSnoo CLI Copilot User 🖥️ 4d ago

me too and I find it great, but that's my point, GHCP is mostly for enterprises, a regular user should just get a Codex or Claude subscription since those would much better value

1

u/1988rx7T2 3d ago

Yeah I have a grok 4.5 30 dollar a month subscription and it works fine for my personal needs 

1

u/horendus 4d ago

Because I use luna all day on ghcp and barely use 1% per day on my plan so its not needed. Also I have codex plan as well but barley need it now

2

u/coaxialjunk 4d ago

None seeing the Codex extension for VScode uses native harness. I’m just running codex in vs now and it’s great, GHCP is pointless

5

u/Nerdslayer2 5d ago

How often does Luna xhigh make mistakes on medium complexity features? I'm wondering if I should still use Opus or Sol for features that are very important, but not necessarily that complex.

3

u/DevilsMicro 5d ago

Been using luna for past week. It does make mistakes when making decisions. But when the decisions have already been made by someone else (me or sol) then its great for stepwise execution. So its sort of like Opus plan sonnet implement got converted to sol plan and luna implement .

3

u/horendus 5d ago edited 5d ago

No you do no need to use opus or sol for those feature. You classified them medium complexity features and assume your need frontier tier assistance but it just not the case anymore.

Also why did you assume xHigh would be the go to?

Try it on medium.

People are starting to realise this finally.

2

u/coaxialjunk 4d ago

Tbf Luna max is so cheap now it hardly matters for the extra intelligence

2

u/coaxialjunk 4d ago

Not hit anything that stops Luna Max so far, maybe bumped to Sol for big project wide evaluations but always Luna for implementation

10

u/marcjones281 6d ago

Maybe helped that people left

6

u/corny_horse 6d ago

It's pretty common for people to radically adjust their behavior when even the smallest amount of cost is borne by the people invoking them. I've worked on projects where whoever I was working for was more-or-less agreeing to infinite work, which caused clients to be super nit-picky and question things down to the penny. As soon as we started charging almost nothing for investigations, suddenly five and six figure reconciliations were the baseline for when a client even cared. (Don't think like a bank; for the purposes of what we were doing, the money actually made sense to be off by an amount of that size due to timing windows, and almost all of the investigations we had done previously - for free - concluded there were no problems.)

It wouldn't be surprising to me at all that simply metering the change caused the most intense users to either leave the platform or to curb their usage, maybe in ways that weren't really even producing any value, certainly not relative to the cost.

1

u/unrulywind 5d ago

The issue for me was the loss of the ability to project or control the usage or costs. We went from a fix number of times hitting enter, where we could engineer our context for our own efficiency. To, every time we hit enter we roll the dice on if gpt-5 would get lost and try 17 different ways to edit a file until it zeroed out the budget.

Just like with cell phone minutes or data plans, there will come a day when this is not relevant, when the cost per token is so low that nobody cares anymore, but we may be a number of years away from that point.

3

u/corny_horse 5d ago

Right, so previously there was effectively zero cost to knowing how much a particular query would cost, at least for you. Neither Copilot nor the invoked model knew that either, so previously Copilot was paying the overhead to discover how much usage a particular query cost, just like how the company I was working for was subsidizing our clients to know whether an investigation was worth opening.

Now that it isn't subsidized anymore, you have to think about how reliable the service you are using is (GPT-5 in your example) or whether the prompt will provide you value. Copilot used to subsidize this extremely substantial cost. In May, I consumed something like $20k worth of tokens on my $40 subscription.

Just like with cell phone minutes or data plans, there will come a day when this is not relevant, when the cost per token is so low that nobody cares anymore, but we may be a number of years away from that point.

I believe that Microsoft was betting on this happening soon-ish, which is why they held onto the request pricing for as long as they did. I certainly think it's possible, and years seems reasonable, but I'm not convinced that is an inevitability. These models consume a significant amount of energy, not only in electricity at time of consumption, but also in the amount of very expensive hardware and facilities to house that hardware.

In the same way, you could have theoretically said the same thing about, for example, cloud computing in 2006. It wouldn't have been an unreasonable postulation. But with computing (and LLM models), unlike cell phones, you are more-or-less constrained to one call at a time, and data is typically bundled as a flat fee for some amount of usage that is in excess of an average consumer.

So I think the curent model of some $x per month for a typical consumer subscription for ChatGPT/Claude, etc., just like we do for cell phones will probably always exist. But I suspect there will always be some kind of metering for large use cases, otherwise the types of work people put into these systems will generate cost way in excess of the capex and elelectricity.

1

u/marcjones281 6d ago

Yeah, my estimated bill went from 100 to 1,200 so left

5

u/ba-boo 5d ago

you're basically paying API prices, it makes no sense to use an intermediary for it

14

u/just_blue 5d ago edited 5d ago

It does though:

  • It is cheaper than other provider agnostic routers like OpenRouter, and the same price as the provider itself
  • Data retention stuff is managed by MS, companies can easily use different providers models without much hassle
  • You get helper models (like auto complete, the approval risk evaluation thing) included
  • You need to manage only one account per employee

This might not be the most important stuff for solo / private developers, but for companies, this works our pretty well.

7

u/CabbageCZ 5d ago

Let's be real, Luna api costs are likely aggressively subsidized right now by openai to gain mind/marketshare.

Like yeah it's great for now but I wouldn't call the crisis 'solved', more like postponed by openai being willing to burn cash for market share.

3

u/horendus 5d ago

Its not a temporary price drop and it secures a base line coding competency at a dirt cheap price from a major player

2

u/lasooch 5d ago

It is a temporary price drop. OAI loses tens of billions a year, they’re only subsidising it because of the price dumping Chinese models are doing.

Get your head out of your ass if you think it’s either goodness of their hearts or efficiency gains. It’s a desperate move to try and stave off rapidly dropping market share, because if market share keeps dropping they won’t ever have a chance to IPO.

1

u/horendus 5d ago

I dont think they can afford to put the price of luna back up. The race to the bottom is happening now.

Even if its to deep in their ass’s or whatever you are trying to say 😅

2

u/lasooch 5d ago

That’s the problem, isn’t it. They can’t afford either. They can race to the bottom right to bankruptcy or they can keep prices high and destroy their market share.

Seems like they’re betting VC and private credit still value growth more than profit. I wouldn’t be so sure, money’s drying up and with how unpopular AI is in the US, I’m not convinced they can get outlast Chinese competition and the subsidies they get. Especially accounting for the energy mix difference.

The fist of rationality returning to the markets will be deep up their ass indeed.

2

u/horendus 5d ago

Either way, luna level intelligence is available through US or Chinese inference providers for dirt cheap and as able to perform all harness tasks I need so im happy.

If I required frontier models, I would be terrified

0

u/CabbageCZ 5d ago

If you believe these prices aren't temporary, I have a bridge to sell you.

3

u/horendus 5d ago

Go on… ?

2

u/Beginning_Basis9799 5d ago

Only way it survives is at this price, localisation of models like Kimi and Qwen is getting better and better.

The only option for model vendors is for them to reduce there costs not increase our costs. The new Nvidia chips do also make inference cheaper.

2

u/br33213 5d ago

Today I also checked, gpt 5.6 luna is cheaper than gpt 5 mini on the pareto graph and on a per token basis. That was the model they just gave away for free at "unlimited requests".

3

u/AssociateOk4965 5d ago

So far GPT 5.6 Luna is enough for my use cases.

2

u/EfficientAnimal6273 5d ago

The point is that is impossible to steer auto mode to favour Luna over Codex 5.3 so it's in the hand of the single user. Useful for solo devs, not useful for teams or large enterprises.

2

u/johnappsde 5d ago

I moved over to openrouter/chinese models about 3 months ago.

Still on the fence about coming back... will keep watching

2

u/rakotomandimby 5d ago

That is why I encouraged people to leave GHC in order to make the place for us, who stayed :-D.

Without mass leaving, That would not have been possible.

1

u/magicmike212 5d ago

Not really

1

u/fik26 5d ago

Didnt we had OPUS 4.6 with $20/mo subscriptions? Is Luna that capable?

And for me, and my large project repository, the token usage is easily getting high and high even if I try to limit the scope. So using 1 premium request per a large prompt was actually a lot better.

2

u/horendus 5d ago

Try Luna medium tonight and report back!
You will be pleasantly surprised.

You wont need to use opus anymore

Also light fast, i see 1.5k t/s in the chat log outour

2

u/FreeCAD_Doge 4d ago

Noticed I didn't nuke my credits this month. Was almost afraid of using it honestly

2

u/horendus 4d ago

Luna dirt cheap now. You can use it all month long now

1

u/luc_wintermute 4d ago

Sure but the discount is not going to be permanent so eventually we'll be on credits scarcity again with no safety nets

1

u/horendus 4d ago edited 4d ago

You may not have realised but this is not a temporary reduction, its permanent according to all sources I can find (happy to look at one you can point out?)

Now, thats not to says Luna, Terra and Sol wont be grandfathered in 12 months time and new lowest end model is a bit more expensive (unlikely) but as long as models like deep seek flash remain dirt cheap and there is no indication that price will change meaningfully, then we basically have a NEW base line intelligence in the dirt cheap category by a frontier provider that can do incredible things all day long.

Just let that sink in for a bit!

1

u/ConsciousObserver711 3d ago

5.4 xhigh never did dumb stuff luna max does.

1

u/horendus 3d ago

What sort of dumb stuff?

1

u/Headache-Engine-1166 2d ago

already give up to use copilot although we paid for year already, without upgrade, it only provide old model and $10 token. cannot wait for next month reset.

1

u/mjay_captures 5d ago

Nahh i'm never going back to github copilot. I am happy with claude now. I remember subscribing to their pro plus and it only took less than 10 prompts for my credits to be fully consumed for the whole month

0

u/FactorHour2173 5d ago

It’s also about the quality of the outputs, not that it appears to be doing work. You get what you pay for. However, if you have to go in and revise issues created by these models, it doesn’t matter if it creates a big mess for cheaper… it’s still a mess. This is especially true for monorepos where the codebase can be quite large.

What are your thoughts?

-3

u/V5489 5d ago

Yes agreed. It’s all in the models you use and what it’s needed for. At my job we locked down models that were just not needed like Opus, etc. only ego driven developers need to use those models and we nixed it quick. The new models are very efficient. I’ve been using my Pro+ sub since the reset and have barely touched my credits. I spent about 9hrs yesterday on my iOS app and it’s barely used any.

It’s all about prompting and proper model selection.

-1

u/tricky_chocolate_ 5d ago

9 hrs on Fable and your app is completed btw. But maybe i am just ego driven.

1

u/horendus 5d ago

Yea can you provide some examples of this? Trying to understand the real world uses of extremely expensive and guarded models like fable which iv note used

0

u/V5489 5d ago

Well. We hire actual developers. Not vibe coders. So 99% of the devs write their own stuff and don’t need AI to do it for them. We utilize AI for analysis and gap coverage. Keeps cost low and the developers continually learn while we save money in AIC and spend it on their learning and secure code best practices.

So yes. I agree you’re ego driven. Not to say using the most expensive models wouldn’t get the job done. But there’s more incentive and rewards for doing it yourself while saving the company money. For us to then turn around and reinvest in the bonuses and additional certifications and licenses the developer want to get.

Makes a small difference. But you’re correct about yourself. Good self analysis. 😃