Google Gemma 4 Runs on AWS SageMaker NVIDIA T4 at 80% of L4 Speed
A developer has successfully deployed Google's Gemma 4 language model on Amazon SageMaker using an NVIDIA T4 GPU, the smallest GPU instance the platform offers. Benchmarks show the T4 decodes at 77–82% the speed of the newer NVIDIA L4, while producing identical outputs. The deployment required a custom patch to work around a shared memory limitation in vLLM 0.30.0 that normally prevents Gemma 4 from running on older Turing-architecture GPUs like the T4. Among the tested model sizes — 2B, 4B, 12B, and 26B — the 12B variant is the largest that fits within the T4's 16 GB of memory. Python-based MCP tools were also developed to simplify management of the vLLM-hosted deployment on SageMaker.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in