Gemini 3.6 Flash's Thinking Dial Can Cut AI Costs 30x With Accuracy Trade-offs
Google's Gemini 3.6 Flash, which became generally available on July 21, 2026, charges users for hidden 'reasoning tokens' in addition to standard output tokens, billed at $7.50 per million. Independent testing conducted on July 24, 2026, found that the model's reasoning effort setting — ranging from minimal to high — can swing costs by up to 30 times on the same task, dropping a 120-word writing task from $0.03316 to $0.00110. Setting reasoning effort to 'minimal' eliminates reasoning tokens entirely and cuts per-call costs by 91–97%, with no noticeable quality loss on retrieval, formatting, and simple factual tasks. However, the minimal setting caused the model to fail all three attempts at a multi-step arithmetic word problem, highlighting a meaningful accuracy risk for complex reasoning tasks. The model's 1-million-token context window and prompt caching behavior were confirmed to match Google's published specifications in testing.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in