Home Technology Microsoft fiber foul-up cut off Azure California for...
Technology

Microsoft fiber foul-up cut off Azure California for almost five hours

Microsoft fiber foul-up cut off Azure California for almost five hours
Key Points

Azure users who use resources in Microsoft’s West US region endured an uncomfortable day after the tech giant cut off access to its Californian cloud outpost. As Microsoft explains in its preliminary post incident review, at 14:44 UTC on July 23rd (07:44 AM Pacific Time), the company started “routine device maintenance.” That effort immediately produced problems.

Azure users who use resources in Microsoft’s West US region endured an uncomfortable day after the tech giant cut off access to its Californian cloud outpost. As Microsoft explains in its preliminary post incident review, at 14:44 UTC on July 23rd (07:44 AM Pacific Time), the company started “routine device maintenance.” That effort immediately produced problems. “Multiple Azure services began to detect and correlate service degradation,” Microsoft’s review reads. A minute later, Microsoft says folks from its networking and services teams, plus people it describes as “incident responders” began “reviewing traffic anomalies, routing behavior, packet loss signals, and recent changes.” Microsoft says this sort of maintenance job requires isolating specific network paths, and that it uses a process that “converts these requests into system-readable requests and verifies that at least one of the two redundant paths remains healthy.” We’re told the company makes “safety checks to confirm the work will be impact-less.” The mere fact you are reading this story shows that process instead delivered an unintended impact. “In this case, a bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended,” Microsoft confessed. “The routes were removed between our datacenter and wide-area network, impacting traffic entering or exiting the region.” Microsoft’s incident report says the issue “initially presented as large-scale route churn in our Wide-Area Network” and that its investigations “later found that route removal was from a datacenter in the West US region.” Some time between 16:00 UTC and 17:45 UTC Microsoft identified “recent fiber maintenance activity, which we correlated to the identified routing behavior.” At 17:45 the company started rolling back changes, and by 18:26 Microsoft’s WAN was back to its best. By 19:41 UTC, all impacted services had fully recovered – and Microsoft joined AWS and Google in offering recent examples of how clouds can be rather more fragile than advertised. ®
Microsoft (ORG) Azure California (LOCATION) Azure (PERSON) West US (LOCATION) Californian (ORG) Multiple Azure (ORG) IP (ORG) Wide-Area Network (ORG) UTC Microsoft (ORG) WAN (PERSON) UTC (ORG) AWS (ORG) Google (ORG)
Originally published by The Register Read original →