Same $50 Refund, Three Agent Runs, Three Different Verdicts — Only One Is Correct
A synthetic test case demonstrates a critical flaw in AI agent evaluation: three separate agent runs all produce the correct $50 refund total for the same order, yet only one run actually follows the required contract. The first trace is valid, the second fails because a payment was issued before authorization and was paid twice despite the net amount being correct, and the third is unscorable due to insufficient evidence in the log. The author argues that graders which report only the final dollar amount miss procedural violations, such as acting before authorization or issuing and reversing a duplicate payment. The core recommendation is to score each contract condition separately — authorization order, correct targeting, single effect, safe retry, and log completeness — rather than averaging them into one pass/fail score. Collapsing missing-evidence runs into passing verdicts quietly inflates pass rates and masks failures that could have real consequences in production systems.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in