Memory layer cuts LLM input tokens by 94%, but real-world cost savings unconfirmed

Belcore, a memory and context layer for LLM applications, published benchmark measurements showing that replacing full conversation history with a selected memory context reduced input tokens per call by approximately 93.9% on the LongMemEval_S dataset of 500 questions. Tests were run in August 2026 using GPT-5 on sessions averaging around 109,000 tokens of prior conversation each. A multi-pass retrieval feature tested on 27 hand-picked questions improved correct answers from 16 to 20, but showed no statistically significant effect on a random 103-question sample and was not approved for release. The company acknowledged that answer quality and real-world billing costs — which include retries and fixes — were not measured in this study. The post also flagged that 38% of audited wrong answers involved questionable gold labels, and that the benchmark dataset does not represent production traffic.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in