Developer runs 170 AI agent goals for $0.49, uncovers 10 bugs unit tests missed
A developer building PlannerCritic, an open-source engine where one LLM writes plans and a second reviews them, ran field tests across 157–170 goals costing as little as $0.30–$0.49 in API fees. The v0.1.0 field test uncovered 10 issues — including harness bugs, design flaws, and prompt gaps — that traditional unit tests failed to catch because 57 of 65 assertion files were silently malformatted. Subsequent releases v0.2.0 and v0.2.1 each recorded zero field-test failures after code review caught all 41 bugs before the LLM ever ran. The project illustrates how a zero-failure result can mean opposite things: in v0.1.0 it signaled a broken harness, while in v0.2.1 it reflected a robust regression gate. The author argues that real-goal field testing against a live LLM surfaces system failures that hand-crafted unit test inputs are structurally unable to anticipate.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in