Developer stress-tests home Kubernetes cluster with chaos engineering, documents failures
A software engineer running a 4-node bare-metal Kubernetes cluster on Talos Linux deliberately introduced failures using Chaos Mesh to test real-world resilience. The homelab setup included a Dell OptiPlex control plane, three Raspberry Pi worker nodes, and tools such as ArgoCD, Longhorn, Cilium, Prometheus, and Grafana — all built on roughly $220 of hardware. Despite appearing stable on paper, with etcd snapshots every six hours and replicated storage volumes, the engineer had never conducted a live failure test. Using Chaos Mesh, a CNCF sandbox tool that injects real failures into Kubernetes clusters, pod-kill experiments were run against a production namespace hosting a PostgreSQL database and several APIs. Initial results showed that while StatefulSets and Deployments recovered within seconds, the exercise was designed to surface deeper weaknesses around storage failover, network partitioning, and actual mean time to recovery.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in