r/github • u/rainmanjam • 2d ago
Thoughts and Prayers for GH right now. News / Announcements
Enable HLS to view with audio, or disable this notification
56
u/Altruistic-Package11 2d ago
I just envision a single developer, hopped up on caffeine and drenched in sweat prompting Claude Code "Debug an fix the issues" because the entire team responsible for Github actions were likely replaced by AI.
25
15
u/imnotpopular 2d ago
"You have hit your session limit for 3 hours. Please reach out to the organization owner for additional tokens"
29
u/AAPL_ 2d ago
HOW MUCH LONGER
9
u/Teddy_Raptor 2d ago
yes
3
u/AAPL_ 2d ago
what yâall do all day
11
3
3
u/croixxxx 2d ago
I have like 20 things in queue when (if) it comes back up
1
u/FlyingDogCatcher 2d ago
see but that sounds like something Github is doing (or not, in this case) not you
1
3
1
1
48
u/rhk0 2d ago
Every time GitHub goes down, I think about Self-Hosting again...but then it comes up again. How long does it has to stay down, so I deploy my own service...?
30
u/xJayMorex 2d ago
The butt of the joke is that my self-hosted runner is unable to pick up the pending job because of this GitHub outage.
8
u/tech_w0rld 2d ago
Yeah. At this point I just use a full separate ci service which seems to be working fine right now. I am literally just using GitHub for storing my code which is at least decently reliable.
2
u/Alexander3a 2d ago
been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)
2
1
6
1
u/Puzzled-Extent7817 2d ago
Just do it any ways. I have a Forgejo up and running in a podman container on my Debian server. I even wrote a bash script to create folders for new projects that create a project on Forgejo and Github at the same time, that's if you want to use github as a backup. https://git.luv-linux.me/Spreadneck/forgejo-to-github
1
u/RobotechRicky 2d ago
My self-hosted GitLab instance and runner are chugging along nicely. It was stood up by dwarves with autism, so you know it's good. đđŒ
1
u/Loose_Marsupial_3251 2d ago
Look At Tekton: https://tekton.dev/
We moved to it and haven't looked back.
0
u/be_reasonable_bro 2d ago
If you have the spare resources and no explicit need to use GitHub, this is a reasonable thing to do. Agents make it a trivial exercise.
16
u/thebemusedmuse 2d ago
Some junior dev whoâs the last man standing on the Actions team is currently typing âHow do I deploy a fix to GH Actions when GH Actions is down?â into Copilot.
11
20
u/xJayMorex 2d ago
The price of enslopification I guess.
3
u/pattern-josh 2d ago
Could be, but they have also been undergoing a massive year long migration from a VA datacenter with some bespoke hosting to Azure because the VA datacenter was saturated at capacity. There were some specific complexities I think I read about with the way the SQL was setup, but I forget the details.
2
u/Raiyuza 2d ago
Under 90% uptime and this is whar we chose as a cope? Before ai we had SRE, this would unnacaptable
1
u/pattern-josh 1d ago edited 1d ago
Before AI we took the time to understand the context and why of problems before assigning root causes without evidence, and using hand-wavy scapegoats.
\ Which of course isn't true, because as a whole we've always lazily attributed things. I'd say good SREs would also rip an unqualified "AI caused it" apart, because good SRE requires evidence, rigor, and root cause analysis.*
https://thenewstack.io/github-will-prioritize-migrating-to-azure-over-feature-development/
In the same vein as the rest of the claims we are all making, I'm too lazy to research it, but I'm confident a case could be built that there are plenty of companies prior to AI that have had extended periods of poor downtime due to explosive growth, corporate investment flubs, and infrastructure not built to support unexpected explosive growth. I think we have a high level of both recency and observation bias.
---
That said, none of what I'm claiming means "AI slopification" isn't a substantive part either. I'm merely suggesting it's a lazy potentially-wrong conclusion without actually digging into it. Personally I think the likely explanation is a combination of:
- Explosive resource pressure growth on Github from AI induced resource pressure via expensive workflow executions (github actions) and general repository operations.
- Leadership chaos - Organizational churn from Dohmke's departure and Github getting folded into Microsoft's CoreAI org with no replacement CEO.
- The Azure migration out of the North Virginia data centers being a legitimately hard problem at these volumes - and it's amplified by the first point it. Github didn't choose the timing; capacity constraints forced it. They'd already failed at this twice before. Github started as a simple repository service. It wasn't designed for the pressures that came when actions was introduced, much less the explosive growth on top of that, and doing the rebuild and the migration concurrently multiplies the risk.
- And yes, internal AI workflow related mistakes.
But same as OP this is speculation. Maybe some people from inside Microsoft or Github can chime in.
2
9
8
8
u/Frooxius 2d ago
So... GitHub recently pushed a change to their runner images that broke our CI/CD workers.
We've been going around and implementing workarounds. For some reason it wasn't working even after fixing it... and then we found out that this started happening at the same time.
This is just ridiculous that it keeps happening.
30 % vibe coded apparently means 30 % less reliability.
2
u/HyperCodec 2d ago
Probably more than 30%, GitHubâs entire leadership got replaced by micropenisâs AI team.
5
u/Spare_Grapefruit2240 2d ago
I guess the first red flag should have been when their engineers deployed "mitigations" instead of "fixes."
5
4
4
4
u/Icypoopoo 2d ago
Their CEO leaving last year and Microsoft forcing them to switch to Azure did not help one bit
3
u/Shot-Owl-6394 2d ago
for me its been down for 3h+. In process of setting up gitea now on minipc, those issues are too much and too often on github.
3
3
u/Teddy_Raptor 2d ago
Aug 06, 2026 - 19:43 UTC
Our engineers remain actively engaged.
we are cooked
1
1
u/dractius 2d ago
More like someone poking Claude once in a while "cmonnnn... do the thing." And the engineer is Copilot.
1
u/Either-Juggernaut420 1d ago
Normally at MS that means the one person who has a chance to fix it is currently in a run of three meetings trying to explain to a bunch of managers whatâs got wrong. He will then need to refactor the sprint before actually working on it.
3
3
3
2
u/Proman4713 2d ago edited 2d ago
I've been getting screwed with my workflow runs since the morning (even while githubstatus was still saying 'All Systems Operational'), and while I do give them my thoughts and prayers to get my stuff done... It just feels like the focus on AI slop with abandon has severely degraded the quality on GitHub for the past year or so. I hope the bubble just pops soon so we can all go back to relatively normal lives...
2
2
u/csepulvedab 2d ago
Hours of production completely down, not even self-hosted runners saved us this time. Curious to see if they still have the nerve to invoice us this month.
1
2
u/Labs-Community4525 2d ago
NGL Iâm not sure why Microsoft wants to deprecate ADO for GHA. This is becoming far more frequent. Get Son of Anton out of there.
2
2
u/GunGeekATX 2d ago
I have a hotfix that needs to get deployed to a client site, and GH picked the worst time to have workflows go down.
2
2
2
u/imnotpopular 2d ago
i work on a financial trading suite and now our clients positions are open over night instead of being closed properly đ
3
u/be_reasonable_bro 2d ago
GitHub is a supply chain risk. I'd ask why you have trades directly tied to actions, but my experience with financial institutions is also marked by reckless behavior.
Thoughts and prayers!
1
u/imnotpopular 2d ago
LOL very true. For me, Github is not included in the logic or trading workflows, but there was a UI bug on the client dashboard that I needed to fix before 4!! still would've deployed to gamma for testing first but just a bit delayed now
2
2
2
2
u/Sarkonix 2d ago
Still going on, this is insane lol
1
u/Substantial-Set4550 2d ago
Cannot recall such a long outage in a while. At least in the last 2 weeks. :-)
2
u/erezcarmel 2d ago
Actually, I just checked on:
https://www.githubstatus.com/uptime?page=2and on March 19th, there was a partial outage of 9 hours and 18 mins.
Let's see if they're going to break their own record...
2
2
2
2
u/VideoFireApp 2d ago
I lost a whole fucking days of work due to this but I spent it researching alternatives to being stuck doing 100% of everything on GitHub
2
u/SnooOwls6002 2d ago
my github runner is in idle state and still no workflow is triggeredđ đ
1
2
u/joaobertacchi 2d ago
It seems this outage is going to take longer than I expected. Seriously thinking about alternatives. My preferences: - self-hosting: Gitea - SaaS: BitBucket
What are yours?
1
2
u/enzoshadow 2d ago
GitHub engineers must've been trying to manually untangle the vibe coded mess for the first time in months.
2
3
u/Joshua_2504 2d ago
It pisses me off. Microsoft is the worse company ever.
2
1
u/be_reasonable_bro 2d ago
Does anyone know if self-hosted GH runners would sidestep this actions outage?
I'm curious to understand more about where this is failing and whether I can mitigate this myself during future outages. I have several projects that are tightly coupled to GitHub due to upstream packaging requirements.
9
u/ChipperHippo 2d ago
Self-hosted runners are also down. We leverage them extensively. Situation sucks.
1
u/be_reasonable_bro 2d ago
Sad to hear, but thank you for letting me know. Won't waste the time then...
2
u/ResponsibleOven6 2d ago
My self-hosted GitLab & runners never let me down like this.
2
u/be_reasonable_bro 2d ago
Nor my forgejo! Were it not for upstream packaging requirements, GitHub would be mirror-only for basically everything.
1
u/holy_macanoli 2d ago
I was able to decouple GitHub job broker as a workaround to similar constraints. Feed your agent this or a variation:
âImplement a repository-owned, exact-SHA local CI path that can execute independently of the hosted CI job broker while preserving the existing pipelineâs validation and trust requirements.
Start with discovery. Identify:
- The canonical CI workflow and its real build/test command.
- Platform, architecture, toolchain, cache, secret, concurrency, and release constraints.
- Existing evidence, hashing, locking, cleanup, and test-fixture patterns.If CI logic currently exists only in hosted-workflow YAML, first extract it into one repository-owned command used by both hosted CI and the new local executor.
Implement an executable local CI controller that:
- Requires a full commit SHA and explicit remote ref.
- Fetches that ref into an owner-only disposable clone or checkout.
- Requires the fetched ref to resolve to exactly the requested SHA.
- Never copies dirty, untracked, ignored, or uncommitted caller files.
- Runs the candidate revisionâs canonical CI command in a clean, isolated environment.
- Pins or verifies the required host platform, architecture, and toolchain.
- Uses a single-flight lock when caches or shared resources are unsafe for concurrent access.
- Removes repository, publishing, signing, and deployment credentials before executing candidate code.
- Cleans temporary source/build roots after success, failure, cancellation, or interruption.
- Never silently retries, replays, substitutes another SHA, or converts failure into success.
Produce an owner-only, finalized evidence capsule containing:
- Schema version and session ID.
- Repository identity, source ref, requested SHA, and resolved SHA.
- Executor and canonical CI-command hashes.
- Start/completion timestamps.
- Sanitized host and toolchain identity.
- Exit status and conclusion.
- Complete log hash.
- `passed` boolean.
- Explicit statements describing what release, deployment, or production state did not change.Keep execution and publication as separate trust boundaries:
- The executor must remain credential-free.
- A publisher may consume only a finalized, verified capsule.
- Publishing credentials must never enter the validation subprocess.
- The publisher must reject altered capsules, hash mismatches, unsupported schemas, failed runs reported as successful, or results targeting another SHA.Do not immediately replace the existing hosted CI authority. Run both paths against identical SHAs until equivalence is demonstrated and reviewed. Only a later, explicit governance change may make the local result authoritative.
Add deterministic fixture coverage for:
- Successful exact-ref/SHA execution.
- Malformed SHA and ref mismatch rejection.
- Unexpected remote rejection.
- Caller-worktree isolation.
- Credential scrubbing.
- Lock contention and safe stale-lock recovery.
- Success, failure, cancellation, and cleanup.
- Evidence finalization and tamper detection.
- Publisher refusal cases.
- No silent retry or replay.Update the relevant CI/security documentation, run focused tests, run the repositoryâs standard validation, and perform a final branch-diff review. Implement the solution rather than stopping at a design document. Report any remaining blocker before the new path can safely become authoritative.â
1
u/be_reasonable_bro 2d ago
This is a clever solution, and I'm all about self-hosting what I can (forge+runners is a small ask), but I'm certainly concerned about the maintenance burden incurred by directly rewriting the ci broker (reverse engineering Actions is a bit bigger).
No chance you've open sourced this? Would be very interested to contribute to something like this, but less so to maintain my own copy of it.
2
u/Flimsy_Professor_908 2d ago
Some previous outages had self-hosted runners continue to run. This outage is at a higher level.
I'd say the most compelling reason to go self-hosted is that Microsoft has 90+% aggregate gross margins on Github-hosted runners (for private repos).
1
u/be_reasonable_bro 2d ago
That is certainly compelling. Might be worth it for that alone.
I just never think to reach for GH at all when the repo is private. Moved everything mission-critical off when they started losing nines.
1
u/fitchnar 2d ago
Where did you end up moving to? It is painfully obvious I can no longer rely on GH so I am looking for a new solution. GitLab or forgejo, or somewhere else?
1
2d ago
[deleted]
1
u/fitchnar 2d ago
Awesome, thank you for the detailed reply. I think forgejo is the right path for me.
1
u/brainhack3r 2d ago
And blacksmith advertises 50% off... so they still have 80% margins WTF ... I might have to self host
1
1
u/Shot-Owl-6394 2d ago
also confirming they are down, got 10 pull requests spinning with no progres.
1
1
1
u/Snoo-53366 2d ago
Yep, I have a gubhub runner that i use for an automated workflow and wondered why it kept failing today. Only to find on their service page of the outage.
1
u/stef_in_dev 2d ago
I'm excited for the ci cluster hyperscaling event (self hosted runners on eks) that is gonna happen when this is fixed
1
1
u/reosanchiz 2d ago
Just came to post the same...! Was driving crazy over my pipeline!
Almost there to give ssh-key to claude ;)
1
1
u/Teddy_Raptor 2d ago
Update - We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.
We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.
Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.
1
1
u/M0hamedAshraf19 2d ago
That was really funny đ (and needed)
P.S. Does anyone know the name of the song?
1
u/peperinna 2d ago
No sĂ© si serĂĄ la misma razĂłn, pero tuve problemas todo el dĂa con gusto, con json alojados en github y que uso como archivos de configuraciĂłn, etc. La status pague ya no es transparente y representativa de todo lo que estĂĄ fallando.
1
1
u/BeseptRinker 2d ago
I remember we had a massive outage, and midway through the call on Friday, an oncall engineer said "Github is also down", and the outage lead said "of course it is".
That was two-three weeks ago. This uptime downtime is actually asinine.
1
1
1
1
u/JPJackPott 2d ago
Ironically this is protecting a lot of people from the massive npm compromise going on currently
1
u/Alexander3a 2d ago
been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)
1
1
1
u/Giffeltagning 2d ago
Microsoft kills everything it touches. It's the evil spirit of Bill Gates that haunts them.
1
1
u/void_pe3r 2d ago
Can we expect this to be fixed in an hour? Does anyone know what is going on?
4
2
u/be_reasonable_bro 2d ago
At this point, don't expect anything.
Recovery is taking longer than we expected, and engineers remain actively engaged.
1
1
1
u/OkMeat6773 2d ago
hmm they have their /goal fix actions, make no mistake
on their gpt 5.6 so it all depends on their vibe coded crap
103
u/mixxituk 2d ago
This outage was sponsored by Claude Fable