Small AI Model Picks Correct Tool 19 in 20 Times Regardless of Phrasing

A two-stage evaluation of AI tool selection found that a small, low-cost model (gpt-5.4-mini) achieved 90–97% accuracy when choosing the correct tool from a shortlist of five, across varying description styles and request phrasings. The test used forced shortlists containing the correct tool alongside its four closest BM25 competitors, ensuring the model had to interpret verb semantics rather than rely on simple topic matching. Unlike the earlier recall stage — where paraphrased requests caused failure rates as high as 5% — selection accuracy remained consistently high even when users paraphrased requests into synonyms. Argument-filling accuracy was slightly lower at 87–97% under lenient scoring, but description verbosity had no meaningful impact on selection performance. The findings suggest that retrieval quality, not model judgment, is the primary bottleneck in end-to-end tool-calling pipelines.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in