AI memory systems cannot reliably detect what they don't know, new benchmark shows
A developer investigating AI memory agent reliability found that the ability to abstain from unanswerable questions is not a single capability but varies with how far removed a question is from its supporting evidence. A new benchmark was built with questions generated at controlled 'excision distances,' revealing that discrimination between answerable and unanswerable queries improves sharply as more context is deleted, from near-chance accuracy at close range to near-perfect at full topic removal. This finding explains why two public benchmarks, LOCOMO and BEAM, appeared to contradict each other: both were measuring different points on the same hidden curve and reporting them as a flat capability score. Further analysis showed that for roughly two-thirds of test questions, deleting the gold evidence caused no change in retrieval scores, meaning the correct answer was never the top-ranked result in the first place. The author concludes that abstention failures in AI memory systems are structural rather than tunable, and that benchmark comparisons are misleading without specifying the difficulty distance of unanswerable questions.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in