Google DeepMind's DiffusionGemma Hits 1,500 Tokens/sec Using Parallel Diffusion Decoding
Google DeepMind released DiffusionGemma this week, an open-weight language model built on Gemma 4 26B that generates text using discrete diffusion rather than traditional left-to-right autoregressive decoding. Instead of producing one token at a time, the model works on a 256-token block and denoises it in roughly 12 parallel steps, yielding around 20 tokens per forward pass. On a single H100 GPU, the model achieves approximately 1,500 output tokens per second, compared to about 303 tokens per second for a comparable Gemma 4 autoregressive setup. However, the speed gains come with capability trade-offs — DiffusionGemma scores lower than the Gemma 4 baseline on benchmarks like AIME 2026 and GPQA Diamond. The throughput advantage also diminishes at batch sizes above roughly 32 concurrent users, suggesting the model is best suited for low-concurrency, latency-sensitive applications such as multi-step AI agent workflows.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in