SShortSingh.
Back to feed

Memory layer cuts LLM input tokens by 94%, but real-world cost savings unconfirmed

0
·1 views

Belcore, a memory and context layer for LLM applications, published benchmark measurements showing that replacing full conversation history with a selected memory context reduced input tokens per call by approximately 93.9% on the LongMemEval_S dataset of 500 questions. Tests were run in August 2026 using GPT-5 on sessions averaging around 109,000 tokens of prior conversation each. A multi-pass retrieval feature tested on 27 hand-picked questions improved correct answers from 16 to 20, but showed no statistically significant effect on a random 103-question sample and was not approved for release. The company acknowledged that answer quality and real-world billing costs — which include retries and fixes — were not measured in this study. The post also flagged that 38% of audited wrong answers involved questionable gold labels, and that the benchmark dataset does not represent production traffic.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Claude AI Agent Autonomously Fixes Bugs, Pushes Code, and Files Merge Requests Overnight

A development team used a scheduled Claude Code agent to automatically process seven bugs from a 34-item backlog overnight, with three fixes merged before the next workday's lunch. The agent operates in a fully autonomous loop, reading task boards via API, writing targeted code fixes, running tests, pushing branches, and creating merge requests on GitLab. The system integrates with a project management tool called FlowBoard, pulling structured bug context such as reproduction steps, error logs, and environment details to guide the AI's fixes. Anthropic has documented Claude Code's ability to navigate codebases, run tests, and produce working patches when given well-structured input. Developers caution that fix quality depends heavily on the quality of the bug report, making structured ticket data a prerequisite for reliable automated output.

0
ProgrammingDEV Community ·

AI Project Management Tools in 2026 Now Track Both Developers and AI Agents

The role of project management tools is evolving as AI agents like Claude Code and Codex increasingly handle development tasks alongside human engineers. Traditional tools built around tracking deadlines and assignments are no longer sufficient when AI agents are also executing work within the same workflow. Modern AI project management platforms are now expected to offer capabilities such as turning rough requirements into tasks, integrating with developer environments, and allowing AI agents to receive, execute, and report on work. Tools like Sharkly are designed specifically as shared work systems where humans and AI agents collaborate within a unified coordination layer. Key features in this new generation of tools include parallel agent task execution, shared project context, human review workflows, and integration with existing development infrastructure.

0
ProgrammingDEV Community ·

Developer builds tic-tac-toe agent using real fruit fly brain wiring data

A software developer has created FlyCNS-TicTacToe, a project that uses publicly available fruit fly connectome data to simulate neural decision-making in a game of tic-tac-toe. The fruit fly connectome is a complete map of the insect's nervous system, detailing every neuron and its connections, which the developer accessed via the neuPrint database. Real synapse counts and predicted neurotransmitter types from the dataset were used to route game-board signals through a subset of the fly's central brain neurons. Because the biological circuit lacks strategic look-ahead, a minimax algorithm guides move selection while the fly circuit is used to break ties among equally strong options. The project is open-source under an MIT licence, with connectome data credited under CC-BY, and the code is available on GitHub.

Memory layer cuts LLM input tokens by 94%, but real-world cost savings unconfirmed · ShortSingh