Claude Opus 5 Scores 79.2% on SWE-bench Pro, a 10-Point Gain Over Opus 4.8
Anthropic released Claude Opus 5 on July 24, and benchmark analysis shows a 10-percentage-point improvement over Opus 4.8 on SWE-bench Pro, rising from 69.2% to 79.2% with no change in per-token pricing. Rival model Fable 5 retains a narrow 1.1-point lead on the same benchmark but costs roughly double per task. Anthropic's official launch communications relied heavily on relative metrics — such as 'three times ARC-AGI-3' and 'more than double Frontier-Bench' — rather than publishing absolute scores, making independent verification difficult. On the older SWE-bench Verified set, Opus 5 achieved 96.0% averaged across five trials, though that benchmark is now considered near-saturated as a discriminator with multiple models scoring above 88%. Internal life sciences benchmarks showed Opus 5 outperforming Opus 4.8 by 10.2 points on organic chemistry and 7.7 points on protein tasks, figures Anthropic stated as point deltas against a named baseline.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in