Google launches cheaper Gemini Flash models to cut AI agent costs at scale
Google unveiled three new Gemini Flash models — 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — designed to run AI agents more affordably at large scale. The flagship 3.6 Flash uses 17% fewer output tokens than its predecessor and up to 65% fewer on the DeepSWE coding benchmark, priced at $7.50 per million output tokens. Despite improved efficiency, Google acknowledges its models still trail top competitors from US and Chinese labs, with the strategy explicitly prioritizing cost and reliability over raw performance. The budget-focused 3.5 Flash-Lite targets high-volume repetitive tasks at just $0.30 per million input tokens, while 3.5 Flash Cyber is a cybersecurity-specialized model restricted to governments and select partners. Meanwhile, the widely anticipated Gemini 3.5 Pro flagship remains in limited partner testing, and pre-training for Gemini 4 has only just begun.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in