Rewriting Tool Descriptions Boosted AI Agent Accuracy from 34% to 100% for $4
An open-source project called Toolmetry found that vague or outdated tool descriptions — not the AI model itself — were causing agents to fail tasks at high rates. By rewriting only the text descriptions that tell agents what each tool does and how to use it, SQLite task success jumped from 34% to 100%, with the entire experiment costing just $4 in API calls. Researchers identified three root causes: overlapping tool descriptions causing wrong tool selection, implied but unnecessary prerequisite steps wasting tokens, and deprecated parameter references leading to confident but incorrect calls. For example, a git server's success rate rose from 75% to 96.7% simply by updating parameter names in its descriptions to match the current API. The findings suggest that tool descriptions function as a contract between an agent and external systems, and ambiguity in that contract reliably produces failures regardless of model quality.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in