Why AI Energy-Per-Query Estimates Vary Wildly — and How to Read Them
Published estimates of the energy consumed by a single AI model query can differ by orders of magnitude, but the gap largely stems from inconsistent system boundaries rather than disputes over physics. Key variables include model size, output length, hardware utilisation, facility overhead (PUE), and whether training or manufacturing costs are factored in. A core formula shows that energy per token equals accelerator power multiplied by PUE, divided by aggregate token throughput — meaning a heavily loaded deployment is significantly more efficient per token than an underused one. Amortised training cost per query is also a function of assumed lifetime query volume, making it a different and incomparable metric to marginal serving energy. Authors stress that any quoted figure should explicitly state which cost boundaries it includes and whether it reflects marginal serving cost or total lifecycle cost.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in