Cosine similarity still outperforms six rival signals in RAG answerability tests
A developer running experiments on the LongMemEval benchmark found that their retrieval system successfully located the correct session 97% of the time, yet the trust layer refused to use that retrieved answer in nearly half of all cases. The core issue is that 'relevance' and 'answerability' are not the same thing — a document can be topically related to a query without actually containing the answer. Six different signals were tested as potential gate mechanisms, including cross-encoder reranking, hybrid retrieval, and entailment-based scoring, but none outperformed the plain cosine similarity already in use. Even a more sophisticated cross-encoder, which reads queries and documents jointly with full attention, scored topically-related but non-answering passages just as highly as genuinely answerable ones. The findings suggest the problem lies not in threshold tuning or signal choice, but in a fundamental architectural gap between measuring relevance and detecting true answerability.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in