Coding Agent Benchmarks Are Misleading Without Retry and Timeout Disclosures
Current coding-agent benchmarks often report a single pass rate without disclosing how many retries, timeouts, or helper calls were used to achieve it, making results difficult to compare fairly. Two separate labs can publish identical pass rates while consuming vastly different amounts of model effort, rendering such figures closer to marketing than rigorous measurement. A proposed methodology addresses this by requiring each benchmark task to carry a hard attempt cap, a wall-clock timeout, and a full retry ledger recorded alongside the score. The framework also introduces task stratification across repair, greenfield, and regression categories to prevent easy tasks from skewing overall results. A reference scorer is outlined that refuses to calculate a pass rate unless every recorded run includes an attempt budget, timeout value, and helper-call count.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in