Google Launches Gemini 3.6 Flash with Lower Costs and Faster Performance

Google released Gemini 3.6 Flash on July 21, 2026, positioning it as its primary developer-focused model for coding, agentic workflows, and multimodal tasks. The new model improves on its predecessor across key benchmarks, with long-context retrieval jumping from 27% to 54% and average task completion time halving from 2.7 to 1.3 minutes. Output token pricing dropped from $9.00 to $7.50 per million, and the model uses roughly 17% fewer output tokens than Gemini 3.5 Flash, making it more cost-efficient in practice. The model features a 1 million-token context window and an updated knowledge cutoff of March 2026, a 14-month improvement over its predecessor. Compared to Claude Sonnet 5, Gemini 3.6 Flash is approximately 25% cheaper and runs at 304 tokens per second versus around 180, offering a compelling trade-off for high-volume and latency-sensitive applications.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in