Google Gemma 4 26B Runs on One TPU v6e with 15x More KV Cache via QAT Build
A developer has published a step-by-step guide to serving Google's quantization-aware-trained (QAT) Gemma 4 26B-A4B model on a single Google Cloud TPU v6e chip using vLLM. The QAT build uses only 17.43 GiB of HBM and delivers 53,888 tokens of KV cache and 1,283 output tokens per second, compared to RedHat's FP8 build which uses 27.99 GiB and yields just 3,456 tokens of KV cache at 668 tokens per second. The improvement was achieved by repacking Google's own unquantized QAT weights into a W4A16 Compressed Tensors format, filling a gap left by Google's official vLLM support, which excludes the 26B model size. Despite the large efficiency gains, both builds scored within one percentage point of each other on a 3,880-record classification benchmark. The resulting checkpoint has been released publicly on Hugging Face, and the repacked model also loads on NVIDIA L4 GPUs without additional patches.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in