Gemini 2.5 Flash Cuts Inference Cost by Half While Boosting Code Accuracy
Google's Gemini 2.5 Flash model has launched at half the price of its predecessor, while delivering notable benchmark improvements, including a 9-point jump in code generation accuracy on the FrontierCode benchmark. The cost reduction applies to the same API surface, requiring no configuration changes for teams already using Flash in production. Separately, Z.ai's GLM 5.2, a 1-million-token open-weights model, is now the default on eve agents and available for free via Vercel's AI Gateway until August 27. On the tooling side, the AI SDK's new harness-acp package implements the Agent Client Protocol, allowing a single adapter to work across multiple ACP-compatible agent runtimes instead of requiring separate integrations for each. Together, these releases reflect a broader industry push toward lower inference costs and standardized multi-agent infrastructure.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in