Google Cloud Brings vLLM TPU Support for Long-Context Embedding Inference
Google Cloud announced native vLLM TPU support for embedding inference on August 26, 2026, aimed at production retrieval workloads rather than chat generation. The implementation targets Qwen3-Embedding-8B and Qwen3-VL-Embedding-8B models, handling text sequences up to 16K tokens and multimodal inputs exceeding 15K tokens. In benchmark testing, a Qwen3-Embedding-8B configuration on TPU Ironwood achieved over 83,996 tokens per second and 5.13 requests per second using bf16 precision and four-way tensor parallelism. Google engineered solutions for TPU-specific challenges including tensor alignment, chunked prefill for memory efficiency, and a hybrid StepPool design to preserve pooling state across chunk boundaries. To ensure cross-hardware consistency, Google validated vector parity with strict cosine-similarity thresholds of at least 0.999 for text and 0.995 for multimodal embeddings.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in