How a Wrong rm -rf Command Wiped GitLab's Production Database in 2017

On January 31, 2017, a GitLab engineer accidentally deleted the PostgreSQL data directory on the primary production server while intending to resync a lagging replica, destroying approximately 300 GB of data within seconds. When the team attempted recovery, all five backup and replication mechanisms in place were found to be non-functional or improperly configured, including nightly pg_dump jobs that had been failing silently due to a PostgreSQL version mismatch. GitLab.com remained offline for roughly 18 hours, with recovery relying on a manual snapshot taken six hours earlier for an unrelated test. The incident resulted in the permanent loss of around 5,000 projects, 5,000 comments, and 700 new user accounts, though Git repositories and wikis were unaffected. GitLab published a detailed postmortem on February 10, 2017, and CEO Sid Sijbrandij issued a public apology, making the incident one of the most widely cited database failure case studies in the industry.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in