AI Agent Runs Cost Far More Than Expected Due to Repeated Input Tokens
A detailed analysis of real-world AI coding-agent sessions reveals that costs are driven overwhelmingly by input tokens, not output, because the full conversation history is resent at every step. A study of approximately 4,300 coding-agent sessions found that the median LLM step carries around 119,000 cached prefix tokens compared to just 875 newly added tokens and 214 output tokens. This means the accumulated context can be roughly 550 times larger than what the model actually writes in a single step. Agentic coding tasks were found to consume around 3,500 times more tokens than single-round code reasoning, with token usage varying up to 30-fold between runs on the same problem. Providers like Anthropic and OpenAI offer cache-read discounts of up to 90%, but costs spike sharply if the prefix changes mid-run, making prompt consistency a key factor in controlling expenses.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in