Peking University Study Shows Compressing AI Agent Actions Hurts Coding Performance
Researchers from Peking University have proposed a new context compression method for long-context coding agents, which typically balloon to 80,000 tokens by turn 25 of a task. The paper, published on arXiv, argues that existing approaches fail because they compress both tool outputs and the agent's own decisions indiscriminately. The team's method, called Latent Observations, Hard Actions (LOHA), preserves the agent's reasoning and recent tool outputs as raw text while compressing only older environmental outputs into soft tokens. A companion training technique, Anchored Context Distillation (ACD), prevents performance degradation by keeping the model's outputs aligned with its original base behavior. Testing on SWE-bench Verified showed token reductions of 43–57% with no loss in task completion rates across two open-source models.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in