Google shifts TPU management to Compute Engine as Cloud TPU API enters maintenance mode

Google has stopped active development of the Cloud TPU API, limiting it to bug fixes and security patches only, with no published sunset date. New TPU hardware generations, beginning with TPU7x (Ironwood), are exclusively supported through Compute Engine or Google Kubernetes Engine. A developer documented migrating a v6e-1 (Trillium) chip running the Gemma model under vLLM from the old queued-resources API to the new Compute Engine instances workflow. The core change involves replacing TPU-specific CLI flags and commands with Compute Engine equivalents, such as swapping accelerator-type for machine-type and using gcloud compute instances instead of gcloud compute tpus. Developers on v5p and v6e can migrate without special access, while TPU7x flex-start provisioning currently requires allowlist approval from Google's account team.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in