Monitoring engines should avoid alerting on new, untested endpoints

The creator of monitoring product Mycellis argues that monitoring systems often alert prematurely on newly configured endpoints. These alerts can occur during normal startup delays, such as DNS propagation or certificate issuance, before the service is fully operational. Existing solutions like Datadog's new_group_delay or Better Stack's 'Pending' status use time-based grace periods to prevent this. The article proposes that systems should instead require a minimum number of successful data points before alerting on failure. This approach ensures alerts reflect actual service issues rather than initial deployment states.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in