IBM Study: AI Agents Passing 77% of Tests May Succeed Only 53% in Real Use
IBM Research has identified a 'consistency gap' in AI agent performance, revealing that a ReAct agent using GPT-4.1 achieved a 77% per-run pass rate on the AppWorld benchmark but completed the same task successfully across all five repeated runs only 53% of the time. The 24-point gap arises because small variations in early steps of an agent's reasoning loop compound across subsequent steps, producing different outcomes for identical inputs — even without any code, prompt, or model changes. To address this, the researchers built a Consistency Analyzer that identifies the exact step where agent trajectories diverge, then generates targeted guidelines stored as episodic memory for future similar tasks. This approach raised the all-five pass rate by 16 points on known tasks and by 13 points on unseen but similar tasks. The findings challenge standard AI testing practices, where each test case is typically run only once, potentially masking significant real-world unreliability.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in