Misconfigured AWS governance tags pulled EKS nodes into Prometheus scrape pool

During an org-wide AWS tagging push on September 8, EKS worker nodes in a production VPC inadvertently received tags that matched Prometheus's EC2 service discovery filters. Because the nodes carried environment=prod and a platform techteam value, Prometheus began attempting to scrape Telegraf metrics on port 9273 — a host agent that was never installed on Kubernetes worker nodes. This caused five EKS targets to register as permanently down, triggering a wave of critical Telegraf Down alerts the following Thursday. What initially appeared to be a fleet-wide monitoring failure turned out to be a tagging side effect from three days earlier, compounded by unrelated exporter issues under the same alert name. The incident highlighted how infrastructure governance changes can silently alter monitoring discovery scope when tag-based scrape filters are not isolated from node management tags.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in