GCP us-west1 Outage: How a Fiber Maintenance Job Took Down 20+ Services

On August 20, 2026, a planned fiber-optic maintenance operation inside Google Cloud's us-west1 region unexpectedly reduced inter-campus network capacity, triggering a two-hour, twenty-two-minute outage. Automated traffic rerouting systems failed to compensate for the capacity loss, causing congestion that cascaded into the region's core infrastructure. The disruption reached critical shared components — including Spanner Paxos consensus and the Unified Metadata Server — which then propagated failures across more than two dozen Google Cloud products, spanning compute, storage, databases, and identity services. Google's incident report highlights that the outage's broad blast radius was shaped by shared underlying infrastructure rather than any direct fault in individual products. The incident underscores that cloud failure domains are defined by shared physical and logical infrastructure, not by the list of services a customer happens to be running.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in