Tollgate v0.2.3 Adds Long-Context Pricing Tiers to Prevent LLM Budget Surprises
Tollgate, an open-source LLM gateway that reserves and settles costs per request, has released version 0.2.3 with improved handling of long-context pricing tiers. Some AI providers, including Google for Gemini 2.5 Pro on Vertex AI, charge a higher rate on the entire request once a prompt crosses a token threshold, not just on the excess tokens. A proxy storing a single flat rate per token class will silently under-charge such requests, potentially causing invoice discrepancies worth thousands of dollars. The update introduces configurable per-model thresholds and rate multiples, applied consistently to both cost reservations and final settlements to prevent drift. Until a tier is explicitly configured, Tollgate logs oversized prompts as possible under-charges rather than applying an assumed uplift.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in