AI Agent Reliability Lags Capability by 18 Months, METR Research Finds
Research organization METR has tracked AI agent task performance since 2019, measuring how long a task—in human-expert time—a frontier AI agent can complete autonomously. Their NeurIPS 2025 paper finds that this capability has been doubling roughly every seven months, but that figure reflects only a 50% task success rate. When measured at an 80% success threshold, agents can handle tasks several times shorter, placing reliable performance about 18 months behind headline capability. METR researchers argue this gap is not primarily a model quality problem but a specification problem—agents fail when tasks contain ambiguity or edge cases, not because the underlying model degrades. The practical implication is that organizations deploying agents should categorize tasks by the cost of a wrong run before scheduling them for unattended operation.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in