Production Kubernetes Outage Reveals Hidden Risks in Deployment Controller Defaults
A three-hour outage on June 17, 2026 took down every WordPress site on a K3s cluster hosted on Hetzner Cloud after a routine node upgrade exposed critical gaps in Kubernetes workload controller behaviour. The incident affected a cluster of 5 nodes running 50 Deployments, 6 DaemonSets, and 12 StatefulSets with 26 persistent volumes. Engineers found that the default RollingUpdate strategy in Deployments causes a Multi-Attach error when a ReadWriteOnce volume is still bound to an old node while a new pod attempts to start on a different one. All 15 single-replica Deployments mounting ReadWriteOnly Hetzner volumes — including WordPress sites, Nextcloud, and Grafana — were identified as vulnerable to this stall condition. A post-incident review led to a revised rule: choose between Recreate and RollingUpdate update strategies deliberately based on storage type, rather than relying on defaults.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in