Two ways a simple LLM token counter goes wrong (with a demo you can run)
A common first version of an LLM spending limit looks like this: let used = 0; async function call(prompt: string) { if (used + ESTIMATE > LIMIT) throw new Error("quota exceeded"); // check const res = await provider.complete(prompt); // call used += res.usage.totalTokens; // add return res; } It reads correctly. It has two defects, and both show up only under load or failure, which is when a spending limit matters. Everything below is reproducible without an API key. The "provider" in the demo is a 5 ms timer. The code is in llm-quota-guard (MIT).
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in