> We installed as much hardware as available power allowed in our existing data centers while accelerating our migration to Azure.
And from the RCA [1]:
> The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic.
"While accelerating our migration to Azure," meaning, they will only solve problems if it helps them also use Azure more.
It is unbelivable that aload of 2.8b commits was totally fine, and a load of 2.9b was a sitewide outage, unless they have no reporting or their tooling is completely incompetent. If things can fall apart so easily, throwing more capacity at the problem won't fix it.