Why Infra Failures Must Be Separated From AI Coding Agent Benchmarks
A new methodology argues that benchmark scores for AI coding agents are often skewed because infrastructure failures — such as dropped SSH connections, full disks, and rate limits — are incorrectly counted as agent errors. The proposed approach requires labeling every failed or incomplete run into one of four categories: MODEL, HARNESS, INFRA, or FLAKY_TEST, before calculating any performance score. Only MODEL-labeled outcomes should factor into an agent's skill score, while INFRA and other non-agent failures are reported separately but excluded from the denominator. The method also recommends maintaining a frozen holdout task suite with clear selection rules, including negative controls and multi-file edits, to prevent inflated results. Researchers are advised to publish split metrics — such as infra censorship rate alongside resolved rate — rather than a single pass rate, so reviewers can audit the integrity of the evaluation.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in