Claude Sonnet 4.5 Beats GPT-5 on Coding Benchmarks, GPT-5 Wins on Price
A head-to-head comparison of Anthropic's Claude Sonnet 4.5 and OpenAI's GPT-5, verified as of September 2026, shows each model leading in distinct categories. Claude Sonnet 4.5 outperforms GPT-5 on coding tasks, scoring 77.2% versus 74.9% on the SWE-bench Verified benchmark, and 50.0% versus 43.8% on multi-step terminal work measured by Terminal-Bench 2. GPT-5 holds a significant cost advantage, priced at roughly 60% less on input tokens and a third less on output tokens per million compared to Sonnet 4.5. GPT-5 also leads on competition mathematics, reaching 94.6% on AIME 2025, and on multimodal understanding with an 84.2% score on MMMU. Developers prioritising code generation and agentic workflows are better served by Sonnet 4.5, while those handling high token volumes, image-heavy inputs, or complex mathematical reasoning may find GPT-5 more cost-effective.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in