Agent cost tracking is broken: token counts don't reveal which AI call spent what

Most AI cost dashboards show correct totals but cannot break down spending by individual agent run, tool call, or retry, making it impossible to identify what drove a bill higher. OpenTelemetry's GenAI conventions currently define no cost attributes in their published spec, and the open proposal has remained unresolved since August 2026. Cloud cost monitors like AWS Anomaly Detection operate on up to a 24-hour delay and evaluate only about three times daily, meaning alerts arrive after money is already spent. Complicating matters further, identical tokens can be billed at vastly different rates depending on cache state, creating up to a 12.5x price swing that raw token counts do not capture. Without per-span cost attribution built into the observability stack, engineers are left unable to trace large bills — illustrated by one reported case where roughly 100 agent instances accumulated over $1.3 million in a single month.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in