High AI agent costs stem from redundant context, not overspending, experts say
Companies running AI coding agents are facing unexpectedly large token bills, with Uber's CTO reportedly stating the firm burned through its entire 2026 AI budget within months. Analysis shows that roughly 73% of costs come from input tokens — meaning companies are repeatedly re-sending the same context to models rather than paying for generated output. Techniques such as prompt caching, model routing by task difficulty, and batch API processing can reduce costs significantly, with a model switch from Claude Opus 5 to Haiku 4.5 alone cutting a sample fleet's monthly bill from $4,125 to $825. However, the most durable savings come from reducing how many tokens are sent in the first place, rather than just lowering their per-unit price. Research using a task-scoped memory layer showed a 41.2% reduction in token usage and 72.6% lower model costs compared to standard retrieval-based approaches.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in