Fable 5.1 Tops Real-SWE Benchmark at 38.8%, but Costs Nearly $7 Per Task
Specific Labs released the Real-SWE benchmark in September 2026, evaluating eight AI model-and-harness combinations across ten tasks drawn from real enterprise codebases covering billing, tax, migration, and infrastructure work. Fable 5.1 paired with Claude Code achieved the highest resolution rate at 38.8% and an estimated cost of $6.96 per rollout, largely due to its high token output averaging over 80,000 tokens on several tasks. GPT-6 Astra via Codex CLI ranked second at 33.8% for $4.67, while Gemini 3.8 Flash via Gemini CLI offered the best value at 31.2% for $2.50 per rollout. Confidence intervals among the top three overlap significantly, meaning the performance differences are directional rather than definitive. Notably, longer agent runtimes did not improve outcomes, and missed requirements — not syntax errors — were identified as the primary failure mode across all tested models.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in