AI Tool Retrieval Fails 95% of the Time When Users Paraphrase Requests

A structured evaluation of deferred AI tool loading found that retrieval works nearly perfectly when users mirror the exact vocabulary in tool descriptions, but collapses to around 5% recall when they paraphrase naturally. The experiment tested BM25 retrieval across 100 synthetic enterprise tools at three description detail levels — terse, realistic, and verbose — using 200 tasks split equally between vocabulary-matching and paraphrase queries. Verbose descriptions, which cost roughly double the tokens of realistic ones, showed no meaningful improvement in paraphrase recall, exposing that adding more words drawn from the same vocabulary provides no extra retrieval surface. Widening the shortlist from 5 to 10 tools only pushed paraphrase recall from 5% to 10%, doubling context costs for minimal gain. The core finding is that a vocabulary gap between how tools are described and how users actually speak creates a near-invisible failure mode, where the model receives a plausible but wrong shortlist and silently improvises.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in