Clock Drift on One Server Derailed a Payment Incident Investigation for Hours
A payment system incident at a software team turned into a two-hour wild goose chase after logs produced a physically impossible timeline, with a gateway appearing to respond before a request was even sent. The root cause was a single host whose clock was running eleven seconds ahead of the rest, a drift that had accumulated over months after a base image rebuild accidentally dropped the time synchronisation daemon. The skewed timestamps also caused an error-rate alert to misplace records across time windows, masking a genuine spike, and explained a certificate validity failure that engineers had previously dismissed as intermittent. In response, the team added clock offset as a monitored metric with alert thresholds, and updated the log pipeline to record both emitted and ingest timestamps so drift becomes immediately visible. They also adopted a broader principle of not inferring event ordering across separate machines from timestamps alone, treating cross-host time comparisons as independent claims rather than ground truth.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in