AKS Supports Near-Bare-Metal Workloads, But Platform Teams Face Key Caveats
Azure Kubernetes Service (AKS) allows teams to use isolated VM SKUs that occupy entire physical hosts, reducing hypervisor overhead for demanding workloads like GPU inference and latency-sensitive financial systems. These specialty SKUs can offer predictable NUMA topology, single-tenant hardware isolation, and full access to features like SR-IOV and DPDK, though compatibility with AKS CNI configurations requires validation. However, using such SKUs does not change core AKS node lifecycle behavior — auto-repair, OS upgrades, and cordon-drain-replace mechanics remain in effect. A key risk is that isolated VM families often have limited regional capacity, meaning surge nodes during upgrades may fail to provision, potentially stalling the process. Platform teams are advised to verify current supported VM sizes for their target region and reassess their upgrade and auto-repair assumptions before committing to specialty node pool designs.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in