Gemma 4 QAT Weights Deliver 2x Faster Decoding Than bf16 on AWS SageMaker L4
A developer benchmarked Google's Gemma 4 E2B model in two formats — full-precision bf16 and a quantization-aware trained (QAT) 4-bit weights checkpoint — on an Amazon SageMaker real-time endpoint powered by a single NVIDIA L4 GPU. The QAT variant decoded tokens at 105.1 tokens per second compared to 51.3 for bf16, and handled 16 parallel requests at 1,077 tokens per second versus 619 for the standard model. Both checkpoints produced identical results across 40 test questions, suggesting no accuracy loss from quantization. The test was conducted on an ml.g6.xlarge instance in the us-east-2 region using AWS's vLLM SageMaker container, with the only configuration change being the model checkpoint variable. A companion suite of Python MCP tools was also built to streamline deployment and management of the vLLM-hosted endpoint.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in