Developer audit finds session history, not AI errors, drives 80% of LLM costs
A developer analyzed 660 AI session transcripts spanning 45 days to identify what was driving up costs in their local AI assistant setup. The audit found 729 failed tool calls, but those recovery costs amounted to only 1.5% of total spending — far less than expected. The dominant cost driver turned out to be context re-reading: as sessions grow longer, the model re-processes its entire conversation history on every reply, silently inflating token usage. Three long-running sessions alone accounted for 20% of one week's total token burn, with the worst single session consuming 240 million cache-read tokens across 665 messages. The developer responded with three practical fixes, the most impactful being a strict habit of archiving long sessions and starting fresh with a brief handoff summary.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in