AI Token Prices Are Falling, But Engineering Teams Are Paying More Than Ever
OpenAI recently cut GPT-5.6 Luna pricing by 80% and GPT-5.6 Terra by 20%, while Claude Opus 5 launched on Amazon Bedrock at competitive rates, continuing a broader trend of falling AI token costs. Despite cheaper tokens, many engineering teams are reporting higher AI-related cloud bills each quarter. The core reason is the Jevons Paradox: lower prices unlock previously unviable workloads, driving far greater usage and ultimately higher spend. Beyond token costs, new layers of AI infrastructure — including agent runtimes, memory stores, observability tooling, and GPU pools — now make up a growing share of total AI spend, estimated to shift from roughly 20% in 2024 to around 70% of the total bill. Unlike token pricing, these infrastructure components follow standard cloud cost curves and require traditional FinOps practices, not vendor price cuts, to control.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in