AI Coding Models Doubled Key Benchmarks in 12 Months, Costs Fell 20-Fold
Between August 2025 and August 2026, frontier AI coding models underwent a dramatic transformation, with SWE-bench Verified scores rising from 49% to 95% and context windows standardizing at 1 million tokens across all major providers. Anthropic's Claude Fable 5, released in June 2026, leads coding benchmarks at 95% SWE-bench Verified, though access remains restricted and it costs $10 per million input tokens. At the other end of the spectrum, DeepSeek V4 Flash and Google's Gemini 3.5 Flash offer competitive coding performance at as little as $0.14–$0.15 per million tokens, roughly one-seventieth the cost of top-tier models. New benchmarks such as Terminal-Bench, MCP Atlas, and OSWorld emerged to measure agentic and computer-use capabilities that did not exist as formal categories a year prior. Despite these gains, engineers note that code review remains a bottleneck and that models still struggle with high-level architectural decisions, excelling at implementation within established patterns rather than choosing which patterns to apply.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in