Engrava Posts Reproducible LongMemEval-S Benchmark Scores for Versions 0.5.0 and 0.6.0
Memory system Engrava published verifiable benchmark results on the 500-question LongMemEval-S test, with version 0.6.0 scoring 81.6% micro in August 2026 and version 0.5.0 scoring 82.4% micro in July 2026. Both runs used the same canonical scorer, standard GPT-4o reader and judge, and a top-k of 20 retrieved turns, with no changes to anything outside the memory layer itself. Unlike many published benchmarks, Engrava released full reproduction artifacts for both runs, including the older, slightly higher-scoring result that predates the current release. The memory pipeline contains no generative language model — ingestion and retrieval rely on deterministic hybrid search over a typed graph, meaning no LLM tokens are consumed on every read or write operation. The four-question gap between the two versions prompted the team to examine run artifacts directly rather than attributing the difference to either regression or statistical noise.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in