Giving AI Coding Agents More Time Does Not Improve Success Rates, Study Finds
A benchmark called Real-SWE tested leading AI coding agents against private enterprise codebases covering billing, tax, and multi-service workflows. The study found that extending agent runtime made virtually no difference — failure rates hovered around 71–73% regardless of whether a task ran under or over 10 minutes. The top-performing system, Fable 5.1 paired with Claude Code, resolved only 38.8% of tasks, while GPT-6 Astra on Codex CLI reached 33.8%. Researchers noted that agents tend to fail not due to insufficient compute time, but because of gaps in contextual understanding or inadequate scaffolding. The study also highlighted that benchmark scores reflect a specific model-plus-harness combination, warning that comparing results across different scaffolding setups leads to misleading conclusions.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in