Same AI Model Can Cost 14x More Depending on the Provider Serving It

A new analysis of 405 paid models on the OpenRouter routing platform found that 182 models are served by two or more providers, with prices varying by a median of 1.87x and as much as 14.47x for the same model. Beyond cost, providers differ on numerical precision — fp4 versus bf16 — and uptime, factors that are not reflected in the model name or API call. Data residency is another hidden variable, as 39% of multi-provider models are served from locations in more than one jurisdiction, which matters for compliance obligations. Researchers also benchmarked latency across 329 models and graded 330 for Korean-language quality, finding that performance does not reliably track release date or parameter count. The findings, published on Hugging Face, argue that selecting a model and selecting an endpoint are two separate decisions, each with meaningful cost and quality implications.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in