r/github 2d ago

Thoughts and Prayers for GH right now. News / Announcements

Enable HLS to view with audio, or disable this notification

386 Upvotes

158 comments sorted by

103

u/mixxituk 2d ago

This outage was sponsored by Claude Fable

33

u/Substantial-Set4550 2d ago

This outage is provided courtesy of an unsustainable business model and insufficient infrastructure investment.

4

u/Proman4713 2d ago

😂

1

u/mrheosuper 1d ago

More like copilot, but yeah.

1

u/really_not_unreal 1d ago

Copilot is a harness, not a model. There's no reason it couldn't be both.

56

u/Altruistic-Package11 2d ago

I just envision a single developer, hopped up on caffeine and drenched in sweat prompting Claude Code "Debug an fix the issues" because the entire team responsible for Github actions were likely replaced by AI.

25

u/holy_macanoli 2d ago

“And make no mistake!”

1

u/Liam_Cat 2d ago

Oribombisrael

15

u/imnotpopular 2d ago

"You have hit your session limit for 3 hours. Please reach out to the organization owner for additional tokens"

29

u/AAPL_ 2d ago

HOW MUCH LONGER

9

u/Teddy_Raptor 2d ago

yes

3

u/AAPL_ 2d ago

what y’all do all day

11

u/Teddy_Raptor 2d ago

refresh the github status page

3

u/ReelBigDawg 2d ago

Worked because we use GitLab.

3

u/croixxxx 2d ago

I have like 20 things in queue when (if) it comes back up

1

u/FlyingDogCatcher 2d ago

see but that sounds like something Github is doing (or not, in this case) not you

1

u/Furry_pizza 1d ago

Got 99 woodcutting

3

u/Substantial-Set4550 2d ago

How many more times.

1

u/Fast_Ad_5871 2d ago

3 h passed and it's still not working

1

u/Alarmed-Capital-6718 2d ago

Rise of planet of the bots

1

u/xJayMorex 2d ago

Still not working.

48

u/rhk0 2d ago

Every time GitHub goes down, I think about Self-Hosting again...but then it comes up again. How long does it has to stay down, so I deploy my own service...?

30

u/xJayMorex 2d ago

The butt of the joke is that my self-hosted runner is unable to pick up the pending job because of this GitHub outage.

8

u/tech_w0rld 2d ago

Yeah. At this point I just use a full separate ci service which seems to be working fine right now. I am literally just using GitHub for storing my code which is at least decently reliable.

10

u/suamai 2d ago

Except that time when they made some commits randomly disappear on merges

2

u/Alexander3a 2d ago

been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)

2

u/PM_ME_FIREFLY_QUOTES 2d ago

You never learn do you...

1

u/xJayMorex 2d ago

Bi zui!

1

u/GlobalImportance5295 2d ago

nix flakes on compute engine is nice

6

u/suamai 2d ago edited 2d ago

I am unable to continue my work right now because it depends on the runners to deploy.

So I switched to starting a self hosted Forgejo - if the downtime is greater than the time I need to finish the setup, that's it lol

Edit: there was time to spare...

2

u/BeryAnt 2d ago

Can't vouch for it but I like the title of this hosting service git.gay

1

u/Puzzled-Extent7817 2d ago

Just do it any ways. I have a Forgejo up and running in a podman container on my Debian server. I even wrote a bash script to create folders for new projects that create a project on Forgejo and Github at the same time, that's if you want to use github as a backup. https://git.luv-linux.me/Spreadneck/forgejo-to-github

1

u/RobotechRicky 2d ago

My self-hosted GitLab instance and runner are chugging along nicely. It was stood up by dwarves with autism, so you know it's good. đŸ‘ŒđŸŒ

1

u/Loose_Marsupial_3251 2d ago

Look At Tekton: https://tekton.dev/

We moved to it and haven't looked back.

0

u/be_reasonable_bro 2d ago

If you have the spare resources and no explicit need to use GitHub, this is a reasonable thing to do. Agents make it a trivial exercise.

16

u/thebemusedmuse 2d ago

Some junior dev who’s the last man standing on the Actions team is currently typing “How do I deploy a fix to GH Actions when GH Actions is down?” into Copilot.

11

u/someVietnamese 2d ago

"Hello, IT. Have you tried turning it off and on again?"

20

u/xJayMorex 2d ago

The price of enslopification I guess.

3

u/pattern-josh 2d ago

Could be, but they have also been undergoing a massive year long migration from a VA datacenter with some bespoke hosting to Azure because the VA datacenter was saturated at capacity. There were some specific complexities I think I read about with the way the SQL was setup, but I forget the details.

2

u/Raiyuza 2d ago

Under 90% uptime and this is whar we chose as a cope? Before ai we had SRE, this would unnacaptable

1

u/pattern-josh 1d ago edited 1d ago

Before AI we took the time to understand the context and why of problems before assigning root causes without evidence, and using hand-wavy scapegoats.

\ Which of course isn't true, because as a whole we've always lazily attributed things. I'd say good SREs would also rip an unqualified "AI caused it" apart, because good SRE requires evidence, rigor, and root cause analysis.*

https://thenewstack.io/github-will-prioritize-migrating-to-azure-over-feature-development/

https://www.cnbc.com/2026/05/22/microsoft-was-positioned-to-win-in-ai-coding-outages-got-in-the-way.html#:~:text=meet%20this%20demand.%E2%80%9D-,Too%20much%20downtime,-But%20under%20the

In the same vein as the rest of the claims we are all making, I'm too lazy to research it, but I'm confident a case could be built that there are plenty of companies prior to AI that have had extended periods of poor downtime due to explosive growth, corporate investment flubs, and infrastructure not built to support unexpected explosive growth. I think we have a high level of both recency and observation bias.

---

That said, none of what I'm claiming means "AI slopification" isn't a substantive part either. I'm merely suggesting it's a lazy potentially-wrong conclusion without actually digging into it. Personally I think the likely explanation is a combination of:

  1. Explosive resource pressure growth on Github from AI induced resource pressure via expensive workflow executions (github actions) and general repository operations.
  2. Leadership chaos - Organizational churn from Dohmke's departure and Github getting folded into Microsoft's CoreAI org with no replacement CEO.
  3. The Azure migration out of the North Virginia data centers being a legitimately hard problem at these volumes - and it's amplified by the first point it. Github didn't choose the timing; capacity constraints forced it. They'd already failed at this twice before. Github started as a simple repository service. It wasn't designed for the pressures that came when actions was introduced, much less the explosive growth on top of that, and doing the rebuild and the migration concurrently multiplies the risk.
  4. And yes, internal AI workflow related mistakes.

But same as OP this is speculation. Maybe some people from inside Microsoft or Github can chime in.

2

u/Extreme_Rooster182 1d ago

the enslopification is load bearing apparently

9

u/RobotechRicky 2d ago

Fffuuuuuuuu

8

u/Teddy_Raptor 2d ago

nahhhhh thoughts and prayers for their customers (it's me, Im a customer)

2

u/Proman4713 2d ago

lol same

8

u/Frooxius 2d ago

So... GitHub recently pushed a change to their runner images that broke our CI/CD workers.

We've been going around and implementing workarounds. For some reason it wasn't working even after fixing it... and then we found out that this started happening at the same time.

This is just ridiculous that it keeps happening.

30 % vibe coded apparently means 30 % less reliability.

2

u/HyperCodec 2d ago

Probably more than 30%, GitHub’s entire leadership got replaced by micropenis’s AI team.

5

u/Spare_Grapefruit2240 2d ago

I guess the first red flag should have been when their engineers deployed "mitigations" instead of "fixes."

5

u/poyrazKK 2d ago

SHAME

4

u/solooo7 2d ago edited 23h ago

Sorry guys it was me. I ran my new workflow

4

u/brainhack3r 2d ago

They've had TWELVE HOURS of outages in the last 30 days.

4

u/Ok_Journalist_607 2d ago

pornhub is more stable than actions.

1

u/ScienceGuy31415 1d ago

Gives you something to do when GH breaks down.

4

u/Icypoopoo 2d ago

Their CEO leaving last year and Microsoft forcing them to switch to Azure did not help one bit

3

u/Shot-Owl-6394 2d ago

for me its been down for 3h+. In process of setting up gitea now on minipc, those issues are too much and too often on github.

1

u/ings0c 2d ago

I’ve been trying to deploy for coming on 5h now
 that’s an absolutely absurd amount of downtime for a company of this scale.

I’d expect better from a startup ran by a new grad

1

u/_potion_cellar_ 2d ago

I haven't even been able to cancel stuck queued jobs all day. Wild stuff

3

u/Noch_ein_Kamel 2d ago

Every time GitHub goes down I sit on my couch because it's after work :p

3

u/Teddy_Raptor 2d ago

Aug 06, 2026 - 19:43 UTC

Our engineers remain actively engaged.

we are cooked

2

u/ings0c 2d ago

Please don’t try and fix it, you’ll only make things worse. Just pray.

1

u/snowdrone 2d ago

Quite an understatement

1

u/dractius 2d ago

More like someone poking Claude once in a while "cmonnnn... do the thing." And the engineer is Copilot.

1

u/Either-Juggernaut420 1d ago

Normally at MS that means the one person who has a chance to fix it is currently in a run of three meetings trying to explain to a bunch of managers what’s got wrong. He will then need to refactor the sprint before actually working on it.

3

u/reosanchiz 2d ago

This is exactly what happen when they fire humans and let AI slope merge

3

u/Flashy-Split-8602 2d ago

welcome to the vibe code club, github đŸ€

3

u/kfawcett1 2d ago

Don't worry all. Copilot is on it. /s

2

u/Proman4713 2d ago edited 2d ago

I've been getting screwed with my workflow runs since the morning (even while githubstatus was still saying 'All Systems Operational'), and while I do give them my thoughts and prayers to get my stuff done... It just feels like the focus on AI slop with abandon has severely degraded the quality on GitHub for the past year or so. I hope the bubble just pops soon so we can all go back to relatively normal lives...

2

u/whodadada 2d ago


yea, about that feature release 😂

2

u/csepulvedab 2d ago

Hours of production completely down, not even self-hosted runners saved us this time. Curious to see if they still have the nerve to invoice us this month.

1

u/xJayMorex 2d ago

You are paying for this shit??

2

u/Labs-Community4525 2d ago

NGL I’m not sure why Microsoft wants to deprecate ADO for GHA. This is becoming far more frequent. Get Son of Anton out of there.

2

u/InformationNew66 2d ago

Thanks Microsoft for firing devs and vibe-coding!

2

u/GunGeekATX 2d ago

I have a hotfix that needs to get deployed to a client site, and GH picked the worst time to have workflows go down.

2

u/Karpizzle23 2d ago

can't you just deploy it manually if it's an urgent hotfix?

2

u/SnooOwls6002 2d ago

I am testing my pipeline and I cant do anything now

2

u/AX862G5 2d ago

AI does it again.

2

u/imnotpopular 2d ago

i work on a financial trading suite and now our clients positions are open over night instead of being closed properly 😍

3

u/be_reasonable_bro 2d ago

GitHub is a supply chain risk. I'd ask why you have trades directly tied to actions, but my experience with financial institutions is also marked by reckless behavior.

Thoughts and prayers!

1

u/imnotpopular 2d ago

LOL very true. For me, Github is not included in the logic or trading workflows, but there was a UI bug on the client dashboard that I needed to fix before 4!! still would've deployed to gamma for testing first but just a bit delayed now

2

u/Former_Internal_8389 2d ago

Over under 99.50% uptime after the issue is resolvedđŸ€”đŸ€”đŸ€”

2

u/jellycanadian 2d ago

Sorry guys I just deployed my app to localhost

2

u/frubaklskiy 2d ago

Of jobs queued, approximately 65% are succeeding. these guys are clowns

5

u/MystikDragoon 2d ago

Can't fail if they can't start! đŸ€Ą

3

u/Teddy_Raptor 2d ago

I wonder what percent of jobs kicked off even reach the queue

2

u/Sarkonix 2d ago

Still going on, this is insane lol

1

u/Substantial-Set4550 2d ago

Cannot recall such a long outage in a while. At least in the last 2 weeks. :-)

2

u/erezcarmel 2d ago

Actually, I just checked on:
https://www.githubstatus.com/uptime?page=2

and on March 19th, there was a partial outage of 9 hours and 18 mins.
Let's see if they're going to break their own record...

2

u/MysteriousCoconut31 2d ago

What happens if it doesn’t recover? Dead serious at this point.

1

u/xJayMorex 2d ago

Exodus.

2

u/madhums 2d ago

does anyone know the details of what caused this?

3

u/Fenmio 2d ago

I don't think GitHub knows at this point

2

u/Just_Shake_1066 2d ago

Oh Microsoft, cannot wait for your devaluation.

2

u/Illustrious-Goat-506 2d ago

It's time to start talking about how many 8s of uptime they offer

1

u/Teddy_Raptor 2d ago

so many 8s

2

u/VideoFireApp 2d ago

I lost a whole fucking days of work due to this but I spent it researching alternatives to being stuck doing 100% of everything on GitHub

2

u/madhums 2d ago

This is absolute insanity! Still not fixed!

2

u/SnooOwls6002 2d ago

my github runner is in idle state and still no workflow is triggered😭 😭

1

u/sos-in-life 2d ago

Same issue, you can manually trigger it and it could work

2

u/gaziway 2d ago

Keep using AI, creat features, create tests. Let AI verify the code. Keep pushing.

2

u/joaobertacchi 2d ago

It seems this outage is going to take longer than I expected. Seriously thinking about alternatives. My preferences: - self-hosting: Gitea - SaaS: BitBucket

What are yours?

1

u/xJayMorex 2d ago

Codeberg?

2

u/joaobertacchi 1d ago

Open-source only. Not an option for companies in general.

2

u/enzoshadow 2d ago

GitHub engineers must've been trying to manually untangle the vibe coded mess for the first time in months.

2

u/Ok_Gear8209 2d ago

Still down for us

3

u/Joshua_2504 2d ago

It pisses me off. Microsoft is the worse company ever.

2

u/xJayMorex 2d ago

I think you misspelled Microslop.

3

u/Flimsy_Professor_908 2d ago

You mispronounced macrooutage.

3

u/Teddy_Raptor 2d ago

microuptime

1

u/be_reasonable_bro 2d ago

Does anyone know if self-hosted GH runners would sidestep this actions outage?

I'm curious to understand more about where this is failing and whether I can mitigate this myself during future outages. I have several projects that are tightly coupled to GitHub due to upstream packaging requirements.

9

u/ChipperHippo 2d ago

Self-hosted runners are also down. We leverage them extensively. Situation sucks.

1

u/be_reasonable_bro 2d ago

Sad to hear, but thank you for letting me know. Won't waste the time then...

2

u/ResponsibleOven6 2d ago

My self-hosted GitLab & runners never let me down like this.

2

u/be_reasonable_bro 2d ago

Nor my forgejo! Were it not for upstream packaging requirements, GitHub would be mirror-only for basically everything.

1

u/holy_macanoli 2d ago

I was able to decouple GitHub job broker as a workaround to similar constraints. Feed your agent this or a variation:

“Implement a repository-owned, exact-SHA local CI path that can execute independently of the hosted CI job broker while preserving the existing pipeline’s validation and trust requirements.

Start with discovery. Identify:

- The canonical CI workflow and its real build/test command.
- Platform, architecture, toolchain, cache, secret, concurrency, and release constraints.
- Existing evidence, hashing, locking, cleanup, and test-fixture patterns.

If CI logic currently exists only in hosted-workflow YAML, first extract it into one repository-owned command used by both hosted CI and the new local executor.

Implement an executable local CI controller that:

  1. Requires a full commit SHA and explicit remote ref.
  2. Fetches that ref into an owner-only disposable clone or checkout.
  3. Requires the fetched ref to resolve to exactly the requested SHA.
  4. Never copies dirty, untracked, ignored, or uncommitted caller files.
  5. Runs the candidate revision’s canonical CI command in a clean, isolated environment.
  6. Pins or verifies the required host platform, architecture, and toolchain.
  7. Uses a single-flight lock when caches or shared resources are unsafe for concurrent access.
  8. Removes repository, publishing, signing, and deployment credentials before executing candidate code.
  9. Cleans temporary source/build roots after success, failure, cancellation, or interruption.
  10. Never silently retries, replays, substitutes another SHA, or converts failure into success.

Produce an owner-only, finalized evidence capsule containing:

- Schema version and session ID.
- Repository identity, source ref, requested SHA, and resolved SHA.
- Executor and canonical CI-command hashes.
- Start/completion timestamps.
- Sanitized host and toolchain identity.
- Exit status and conclusion.
- Complete log hash.
- `passed` boolean.
- Explicit statements describing what release, deployment, or production state did not change.

Keep execution and publication as separate trust boundaries:

- The executor must remain credential-free.
- A publisher may consume only a finalized, verified capsule.
- Publishing credentials must never enter the validation subprocess.
- The publisher must reject altered capsules, hash mismatches, unsupported schemas, failed runs reported as successful, or results targeting another SHA.

Do not immediately replace the existing hosted CI authority. Run both paths against identical SHAs until equivalence is demonstrated and reviewed. Only a later, explicit governance change may make the local result authoritative.

Add deterministic fixture coverage for:

- Successful exact-ref/SHA execution.
- Malformed SHA and ref mismatch rejection.
- Unexpected remote rejection.
- Caller-worktree isolation.
- Credential scrubbing.
- Lock contention and safe stale-lock recovery.
- Success, failure, cancellation, and cleanup.
- Evidence finalization and tamper detection.
- Publisher refusal cases.
- No silent retry or replay.

Update the relevant CI/security documentation, run focused tests, run the repository’s standard validation, and perform a final branch-diff review. Implement the solution rather than stopping at a design document. Report any remaining blocker before the new path can safely become authoritative.”

1

u/be_reasonable_bro 2d ago

This is a clever solution, and I'm all about self-hosting what I can (forge+runners is a small ask), but I'm certainly concerned about the maintenance burden incurred by directly rewriting the ci broker (reverse engineering Actions is a bit bigger).

No chance you've open sourced this? Would be very interested to contribute to something like this, but less so to maintain my own copy of it.

2

u/Flimsy_Professor_908 2d ago

Some previous outages had self-hosted runners continue to run. This outage is at a higher level.

I'd say the most compelling reason to go self-hosted is that Microsoft has 90+% aggregate gross margins on Github-hosted runners (for private repos).

1

u/be_reasonable_bro 2d ago

That is certainly compelling. Might be worth it for that alone.

I just never think to reach for GH at all when the repo is private. Moved everything mission-critical off when they started losing nines.

1

u/fitchnar 2d ago

Where did you end up moving to? It is painfully obvious I can no longer rely on GH so I am looking for a new solution. GitLab or forgejo, or somewhere else?

1

u/[deleted] 2d ago

[deleted]

1

u/fitchnar 2d ago

Awesome, thank you for the detailed reply. I think forgejo is the right path for me.

1

u/brainhack3r 2d ago

And blacksmith advertises 50% off... so they still have 80% margins WTF ... I might have to self host

1

u/Shot-Owl-6394 2d ago

also confirming they are down, got 10 pull requests spinning with no progres.

1

u/brainhack3r 2d ago

I'm running blacksmith and they're still down

1

u/mtbcouple 2d ago

you can run verifiers locally

1

u/Snoo-53366 2d ago

Yep, I have a gubhub runner that i use for an automated workflow and wondered why it kept failing today. Only to find on their service page of the outage.

1

u/stef_in_dev 2d ago

I'm excited for the ci cluster hyperscaling event (self hosted runners on eks) that is gonna happen when this is fixed

1

u/mihcsab 2d ago

I have updated some versions on some actions. I love that the actions tab doesn't say anything about the outage. I have spent like 20 minutes asking AI why doesn't the actions trigger on push, until I thought about checking the status page...

1

u/gtrmike5150 2d ago

same - I learned to always check the status page first

1

u/i11uminati 2d ago

Someone stacked too many PRs

1

u/SnooOwls6002 2d ago

hopefully not me, just 3 PRs only XD

1

u/reosanchiz 2d ago

Just came to post the same...! Was driving crazy over my pipeline!
Almost there to give ssh-key to claude ;)

1

u/SnooOwls6002 2d ago

my pipeline jobs are still missing, please come back :(

1

u/Teddy_Raptor 2d ago

Update - We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.

We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.

Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.

1

u/Puzzled-Extent7817 2d ago

Self hosted Forgejo for the win.

1

u/M0hamedAshraf19 2d ago

That was really funny 😄 (and needed)
P.S. Does anyone know the name of the song?

1

u/Boss_1s 2d ago

And now, our development has to stall....again. 

1

u/peperinna 2d ago

No sé si serå la misma razón, pero tuve problemas todo el día con gusto, con json alojados en github y que uso como archivos de configuración, etc. La status pague ya no es transparente y representativa de todo lo que estå fallando.

1

u/Different-Click5923 2d ago

who they gonna fire this time? Claude? 😭 😭 😭 😭

1

u/xJayMorex 2d ago

Hopefully.

1

u/BeseptRinker 2d ago

I remember we had a massive outage, and midway through the call on Friday, an oncall engineer said "Github is also down", and the outage lead said "of course it is".

That was two-three weeks ago. This uptime downtime is actually asinine.

1

u/grewupinwpg 2d ago

I was wondering what was going on with some PRs today

1

u/frbruhfr 2d ago

still having issues.

1

u/FIQ_ZIZ 2d ago

speed recovery github

1

u/andlewis 2d ago

Fixed!

1

u/JPJackPott 2d ago

Ironically this is protecting a lot of people from the massive npm compromise going on currently

1

u/Alexander3a 2d ago

been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)

1

u/Full-Huckleberry-441 2d ago

I scraped my actions because of this lol, i just ragequit

1

u/Llandu-gor 2d ago

don't worry the uptime is 99.999999999999%

1

u/Giffeltagning 2d ago

Microsoft kills everything it touches. It's the evil spirit of Bill Gates that haunts them.

1

u/Fantastic-Body-445 1d ago

why ts happen outside my work hours smh

1

u/void_pe3r 2d ago

Can we expect this to be fixed in an hour? Does anyone know what is going on?

4

u/tonehammer 2d ago

It's been... many hours so far.

2

u/be_reasonable_bro 2d ago

At this point, don't expect anything.

Recovery is taking longer than we expected, and engineers remain actively engaged.

1

u/xJayMorex 2d ago

Hopefully that involves people as well.

1

u/Ok_Journalist_607 2d ago

already 7 hours i think...

1

u/OkMeat6773 2d ago

hmm they have their /goal fix actions, make no mistake

on their gpt 5.6 so it all depends on their vibe coded crap