Study finds AI assistants cite wildly different sources for same search queries
A developer set out to measure whether AI assistants like ChatGPT consistently recommend the same businesses when asked buyer-intent questions, only to discover that results varied dramatically across repeated queries. To separate genuine model differences from random noise, the researcher ran 10 local-service questions three times each across four AI assistants — GPT-4o, Claude Haiku 4.5, Gemini 2.5 Flash, and Perplexity Sonar — all with web search enabled. Jaccard similarity scores were used to compare both cross-model consistency and each model's self-consistency across repeated runs. The methodology revealed that simply asking an AI a question once and recording its answer is statistically unreliable, since LLMs sample from a distribution rather than returning a fixed result. The findings highlight a broader challenge for anyone testing LLM-powered features, where traditional input-output testing assumptions break down entirely.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in