Cheap AI Models Now Cost 90% Less, Reshaping Economics of AI Development
A new generation of small, efficient AI models has sharply cut the cost of running AI workloads, with models like GPT-5.6 Luna priced at just $0.20 per million input tokens — an 80% reduction from its predecessor. The cost drop is largely driven by Mixture-of-Experts (MoE) architecture, which activates only a fraction of total parameters during inference, delivering near-frontier quality at a fraction of the compute cost. Models from OpenAI, Alibaba, Zhipu AI, and DeepSeek all now ship at or below $1.20 per million output tokens, compared to $10 or more for top-tier models. The capability gap between cheap and premium models has narrowed significantly, with one production user reporting GPT-5.6 Luna achieved six times lower cost with only marginal quality loss on a cybersecurity benchmark. For developers, the practical implication is that AI-powered applications previously unviable on cost grounds — such as personalized content generation — may now be economically sustainable.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in