AI Memory Benchmark Reveals Grader Errors, Not Model Failures, as Bigger Risk
A team built a 30-question exam to test whether an AI system could accurately recall 130 days of its own operational history. On the first attempt, their own stack scored 11 out of 18, but a follow-up audit found all seven failures were caused by errors in the grading system, including two incorrect answer keys. Testing the same model with different memory configurations showed that disabling write-time semantic compression in mem0 lifted the score from 11/18 to a perfect 18/18, highlighting how a single configuration flag can significantly impact results. The team also re-examined a widely cited cost-efficiency claim, finding that the "444x cheaper" figure only holds when compared against heavy models like Claude Opus, while comparisons with fast models yield a far narrower 16–18x range. To improve transparency, the team open-sourced their grader and introduced signed results, three-state verdicts, and a 90-day window for challenging outcomes.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in