Google Launches Gemini 3.6 Flash Claiming 65% Cost Cut on Long AI Agent Tasks
Google DeepMind has announced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, with Gemini 3.6 Flash priced at $1.50 per million input tokens and $7.50 per million output tokens via its API. The company claims the new model can reduce token costs by up to 65% on long-horizon engineering tasks, where agents accumulate large amounts of context across many steps. Such tasks grow expensive because models repeatedly process prior conversation history, source files, tool outputs, and instructions with each step. However, the actual savings will vary depending on factors like retry frequency, context-caching usage, and whether the model completes tasks in fewer or more steps than its predecessor. Google also confirmed that a Gemini 3.5 Pro model is forthcoming, intended for more complex reasoning tasks that require a higher compute budget.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in