How a Single Buggy Audit Tool Took Down Meta's Entire Network for Six Hours
On 4 October 2021, Facebook, Instagram, and WhatsApp became unreachable worldwide for roughly six hours after a routine backbone maintenance command severed all connections between Meta's data centres. A built-in audit tool designed to prevent such a command from executing had an undetected bug, allowing the instruction to run unchecked. Because Meta's edge DNS servers are designed to withdraw their BGP routes when they lose contact with data centres, every server did so simultaneously, making Meta's services invisible to the internet. Recovery was severely delayed because remote access tools and internal diagnostics both depended on the very network and DNS infrastructure that had gone down, forcing engineers to travel physically to data centres. The outage was ultimately caused by a circular dependency in Meta's recovery architecture and an automated safety mechanism that lacked a fail-safe against a total simultaneous failure across all locations.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in