Metrics Failure Count Alert Polling (with Cron Workers and Durable State)
TL;DR: Count terminal notification delivery failures, query a recent window on a schedule, and notify only when a threshold crossing changes durable alert state. The metric answers an aggregate question; a Postgres row answers the equally important question of whether this worker has already announced that condition. Keep those responsibilities separate, and use a heartbeat monitor for the different failure in which the polling job never runs. For a marketplace notification service, my decision rule is concrete: use a counter-and-poller design when the operational question is “did at least N d
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in