Developer builds LLM gateway to track AI token costs per product after billing blind spots
A developer running multiple AI-powered products — including a thesis search engine, an image processor, and a video tool — found it impossible to tell which product was driving up shared LLM bills. To solve this, they built a gateway proxy that sits between all products and OpenRouter, an OpenAI-compatible model aggregator, logging token usage and cost per tenant with each request. The gateway also adds semantic caching to avoid redundant model calls and enforces per-product consumption quotas. While preparing to publish the project, the developer discovered the README falsely claimed an automatic failover feature that had never been implemented, and corrected this by moving it to an honest roadmap section. The gateway is now live in production, with the thesis search tool as its first connected product, and the full code including architectural decision records is available on GitHub.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in