Terminal-Bench 4.0: Top AI Models Differ by 0.3% in Score but Nearly 2x in Cost

The Terminal-Bench 4.0 leaderboard, published in early September 2026, shows GPT-6 Astra via Codex leading at 58.2% and Claude Fable 5.1 via Claude Code close behind at 57.9%. Despite the near-identical scores, a full benchmark run costs roughly $3,300 for Astra compared to approximately $6,200 for Fable 5.1. The margin between the two models falls within their overlapping statistical uncertainty ranges, meaning neither can be declared a clear winner. At the other extreme, Sonnet 5 scored just 12.4% while costing $9,600 per run — nearly three times the price of the top model for a fraction of the performance. Analysts suggest developers should weigh cost-per-point and score uncertainty, not just leaderboard rank, when selecting a model for production use.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in