How a 6-Hour Silent Webhook Failure Led One Team to Rebuild Its Delivery System
A software team discovered their webhook delivery system was broken only after a customer reported a missing payment status update, by which time over six hours of events had been silently lost with no retries. The root cause was a TLS certificate rotation by an integration partner that triggered delivery failures the system logged quietly and ignored, as no retry logic had ever been built. The team replaced their fire-and-forget approach with exponential backoff retries starting at 10 seconds, capped at 30 minutes with randomised jitter, followed by routing exhausted events to a dead-letter queue instead of discarding them. They also built a replay tool alongside the queue and added alerting when an endpoint accumulates more than 20 dead-lettered events per hour. To prevent duplicate-delivery issues during retries, every outbound event was assigned a UUID and partners were formally asked to implement deduplication on their end.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in