How a stale alert threshold let 600 errors per minute go unnoticed across 34 pods
A checkout service at an unnamed company silently experienced roughly 600 upstream errors per minute for two hours in June after a partner tax API began failing one in eight calls. An existing alert rule, written three years earlier when the service ran on just four large pods, was set to fire only if a single pod exceeded 40 upstream errors per minute. Following a cost-cutting migration to smaller, more numerous pods managed by an autoscaler, the same error load was now spread across 34 pods, keeping each one well below the trigger threshold. Support teams flagged the incident before the alert ever fired, exposing how the autoscaler had silently changed the assumptions baked into the old threshold. The team responded by auditing all alert rules, replacing absolute per-instance counts with service-wide ratio checks and adding a monthly job to flag any rule whose assumed instance count diverges significantly from the live fleet size.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in