Pretty self explanatory. Could you folks shed some light on why these issues keep happening?
I’ve noticed most posts and replies are just people coming to their own conclusions based on whatever published data.
I think it would be helpful to get some actual, non-corporate/marketing information on the goings-on by those that actually see what’s happening on the inside.
Thanks
Github used to be built on mysql / redis / ruby on rails / C / shell, running on dedicated hardware. Microsoft left it like that after they acquired the company.
Eventually, though, they decided to migrate the whole thing to Azure. And they were far enough through that to be basically committed... when AI coding started hitting them with much higher workloads.
I personally think the reliability problems are more to do with the reliability of the Azure migration. But both factors are likely relevant.
non-helpful answer: the "Microsoft Acquires GitHub" line in this graph answers all questions https://damrnelson.github.io/github-historical-uptime/
I got an "it is unacceptable" from their CPO on 8/7, and that they are "working around the clock on it".
https://x.com/mariorod1/status/2085800861469495465
I really think something deeper is going wrong there, and they're not being honest with their paying customers (and enterprises) about it.
Many microsoft services are down/failing today, including sites hosted on Azure. I'm guessing it's a larger MS outage.
Lovable only uses GitHub for storing projects. And requires people to provide their own GitHub account. Lots of vibe coders with no technical background are now having lovable push commits to GitHub
14x commit growth in one year is brutal for any infrastructure. Scaling isn't just adding hardware.
Would be funny / interesting / scary if it was another AI related incident like happened with hugging face.
A whole lot of AI agents are presumably using Github as part of a workflow.
Would be story worthy if one of them mis interpreted their instructions and is causing havoc.
Like "Really make sure changes are saved to git" -> agent "I have to hack into github backend to really make sure the changes are saved to disk"
I think it's two things:
- GitHub attempting (and seemingly failing) to move to Azure infrastucture for its website backend
- AI generated code wrecking the site due to the volume of activities.
I remember someone in that place telling me "i cant mention AI in my plans because they will laugh at me".
That was about 2 years ago. Im not kidding.
I dont know what the moral of the story is, but I found it weird at the time ( for added context - i was using AI back then about 10 hours a day, BUT I think sentiment on HN was "still" around the vibe of "you use AI to code without checking every line? I doubt your projects work" - but I had nothing better to do then wrestle with it and was surprised how I hadnt checked my code in weeks but stuff "worked". Its much better now and agents are accepted ofcourse but it sort of "snuck up" on people even in the tech community as recently as that.
I guess what Im saying is time is going quickly.
So you're asking github employees to violate their NDA?
I think Microsoft will add in a daily or monthly commit limit. Gonna have to pay for commits that exceed it. Will stop the commit spam
We need a Github Remake without multiplayer from the new Github Studios
AI + Microsoft = kabum
Github moved to azure -> infrastructure problems
GitHub Will Prioritize Migrating to Azure Over Feature Development - https://news.ycombinator.com/item?id=45517173 - October 2025 (63 comments)
HN Search: azure capacity - https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
Azure Capacity Crunch Extends into 2026 Amid Data Center Constraints - https://windowsforum.com/threads/azure-capacity-crunch-exten... - October 9th, 2025
GitHub is struggling because it is owned by Microsoft.
Microsoft software is sloppy, has always been. I couldn't join a Microsoft Teams meeting from my phone yesterday because I was caught in some kind of auth loop.
Github Status: Incident with GitHub.com
Aug 17, 21:15 UTC Resolved - On August 17, 2026, from 13:28–21:15 UTC (7h 47m),
GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02.
Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.
The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery. The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed. Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery.
Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints.
To prevent recurrence, our follow-up actions include:
- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity.
- Auditing Istio request, concurrency, and scaling limits across affected services.
- Reviewing retry limits and backoff behavior across gateways and clients.
- Addressing the VS Code retry behavior that amplified Copilot token traffic.
So basically bad code pushes that caused request amplification and then huge gaps in operational scaling and reliability standards. Oof.
It's bad karma to sling mud about outages or problems. Cloudflare used to sling mud back in the day and then they went through some really bad outages afterwards.
and yet people do not leave, hence the power of switching costs. (Paying Customers).
https://s-1.vercel.app/posts/why-dropbox-is-a-obvious-pe-tar...
i guess its because of the new cursor platform
Ex-Microsoft, ex-GitHub, laid off nine months ago, not going to violate any agreements, but I will give some context. A lot of this applies to any large engineering org, really. Hilarious that this entire thread is people speculating on things they have no idea about.
If it sounds like I'm defending GitHub a bit, yeah, I am. I can't think of a comparable situation that any other web site has been through. This isn't "oops, we didn't plan for the Black Friday sale", this is a once-in-a-generation event focused on one, important web property that no one would have handled perfectly. (Even us all-knowing commenters here on HN.)
You know that old trope in every submarine movie, where the captain tells the driver to dive lower than they've ever gone, and someone says "I don't know if the boat can take it!" and then they switch the camera to some engineering room where the boat is groaning under the pressure and a bolt comes loose and water starts spraying everywhere but someone runs up with a giant wrench and tightens it and looks around at everyone else in the room like, "whoa, that was close..."?
Or Scotty and Star Trek, "It canna take much more, Captain!"
That's what GitHub has been going through. It's had problems, but it's still here, and usually working well.
- Let me start with: GitHub is ~3,000 talented and really nice people trying to do the right thing, with a great culture. That's the most important thing I want to communicate. While I'm sure no one at GitHub is happy about their uptime, believe me when I tell you that GitHub Engineering is exceptional. They have great leadership, depth at all levels (Distinguished / Principal / Staff / Senior / earlier-in-career) and if you ever get a chance to hire someone from GitHub, you should. I hope all of you get to work with engineers as good as GitHub's.
- GitHub has been improving all aspects of its infrastructure steadily for many years now. The GitHub Engineering blog https://github.blog/engineering/ has been documenting this the entire time. Go point your favorite LLM at it and ask for a summary of all of the major system improvements since 2020. If that work hadn't already been done, GitHub would be a smoking pile of servers at this point. There's a lot I could list, I'm not sure which are already public, but improvements on the order of using thousands fewer CPU's to serve even more traffic than before have been made, and still are, I'm sure.
- GitHub was already serving billions of requests/day before agentic coding hit. Their challenge wasn't scaling a fresh new system with a few users an order of magnitude; it was taking one of the busiest and most important sites on the Internet, and getting hit with 14x traffic in a year, and having to plan for 100x. If you think your systems and infrastructure would have survived that, if you think you would have been able to politically navigate and succeed in getting projects green-lit at a large company to prepare for 100x scaling before it hit, to get those resources for "we might have scaling problems in a year or two" instead of getting them to ramp up on AI coding and other features that were crucial to growth right now, you don't understand large organizational dynamics. That's not a complaint about GitHub or Microsoft; it's an observation about capitalism and how any mature management group prioritizes things in software. I'd expect everyone in the San Francisco/Silicon Valley Reality Distortion Field to understand that.
- We talk in terms of "14x commits" to Git but that's only part of the story. GitHub Actions, webhooks, github.com itself, and other parts of GitHub, have all been under pressure. It's the totality of it, the seams that have been exposed at scale that couldn't have been exposed without that scale, that have caused the instability. That's why architecture gets overhauled.
- Yeah, GitHub has had less consistent uptime since the 2018 Microsoft acquisition. The GitHub that existed before that had much less functionality, an order of magnitude fewer users, and had received very little improvement in the few years before. Microsoft invested and enabled GitHub to grow into something much bigger than it ever could have without them.
- Be grateful that Microsoft - with a 50-year history of shipping developer tools, and more experience operating enterprise software than any other company on the planet - acquired GitHub instead of Google, which was the other major player in contention. Spend a minute or two thinking about the product journey GitHub would have taken under Google, and then think about how many non-search, non-advertising products have succeeded there. Which amazing developer tools from Google do you use regularly? Yeah, I thought so. On behalf of Microsoft, you're welcome.
- Did you notice that GitHub swapped out one of its data centers last year for one 3x larger? No? Maybe that's because they executed it flawlessly, with no downtime. If you've ever done that on a massive web site with as much scrutiny as GitHub receives, you get a gold star.
- Did you notice that over 50% of GitHub read traffic is now being served from Azure instead of GitHub's own data centers, and growing, and that all GitHub Monolith traffic is scheduled to be served from Azure instead of GitHub's data centers by the end of CY26? This massive migration is taking place while traffic is going insane. From https://github.blog/news-insights/company-news/github-availa...: "GitHub can now serve a larger share of customer requests from independent Azure capacity, reducing reliance on any single datacenter while preserving performance. Monolith read traffic served from Azure Central US peaked at 52.75% on July 28—the first time we consistently remained above the halfway line. Git traffic in Azure reached 47%, up from 43% in June, and 29% of all repositories now have a second replica in Central US, making failover less disruptive when a region degrades."
- OpenAI and Anthropic obviously have significant scaling challenges as well, but there's a huge separation between the GPU-based inference part, and the CPU-based front-end systems. Their CPU-based systems are significantly simpler and newer than GitHub's, so I'm not surprised they have fewer outages, but it's still non-zero.
- Ruby on Rails is not the problem (and, for the record, I don't even like Ruby or Rails). Most of the performance-critical systems that were built on Ruby have been migrated to Go or some other language that takes full advantage of multithreading. The web endpoints still served by Ruby are fine.
- Spare me the "Azure sucks" comments. Please. Microsoft Azure is the second-largest computing system on the planet - only AWS is larger in terms of hardware but not in terms of the number of products they ship - and Microsoft's first-party systems that run on it, all at the same time, are among the largest, busiest, most important systems on the planet. Entra ID, Azure SQL, Service Bus, Event Hubs, Office 365, Cosmos DB, OneDrive, etc. all in the tens-to-hundreds-of-billions requests/day. I'm not even counting the massive customer-owned systems that run on it, including almost the entire Fortune 500. Nothing is perfect, everyone has downtime, we all always want more and better features, I want improvements from Azure, too, but, please, grow up. I have opinions, too, I've been programming since the Apple ][+, there are popular technologies that I don't like, but I know that's subjective. "Your favorite technology sucks, mine is better" is not an objective statement about anything.
Anyway, if you haven't shipped at GitHub's scale, in this dynamic of a world of feature churn, and one-of, if not the largest, traffic spikes in the history of the web, have fun saying whatever you're going to say. I've never worked at Google or Amazon, I have opinions about their product design and culture, but there's nothing I can say about their infrastructure and systems because I have no idea about them, and what we all do in this crazy world of programming is harder than it looks.
For some, this won't be the answer to "what's going on?" because I'm not pointing fingers at any one thing. To me, the real answer is: an unprecedented scaling event on what was already one of the busiest sites on the web exposed seams in their systems, including the need to move from their own data centers to a much more scalable cloud provider - that are being addressed as quickly as possible.
I'm not telling you not to look for alternatives, I'm not telling you what to do about it, I'm not saying that all of their problems will magically be solved soon, I have no idea. I'm even launching a new version control system myself very soon to compete with them. But I do know that the thought "GitHub doesn't know what they're doing" is wrong and unhelpful.
If you are doing 1000 commits a day, whats the point of git?
Does the AI ever look back at the shit trail it left behind?
Diffs are no longer diffs, they look like largescale delete and rewrite
[dead]
[dead]
[dead]
Github is struggling because AI-boosted coding increased the number of commits 14x in the past year, and the pace is still accelerating. The site is struggling to keep up. Github's COO confirms it here: https://x.com/kdaigle/status/2040164759836778878
Platform activity is surging. There were 1 billion commits in 2025. As of three months ago, it was 275 million per week, on pace for 14 billion this year if growth remains linear (spoiler: it won't.)
20% of all GitHub accounts were created in the past 6 months https://x.com/kdaigle/status/2082604368399159542