Mercury 2.5 Tops LLM Speed Charts at 1,107 Tokens Per Second in September 2026
Inception Labs' Mercury 2.5 has emerged as the fastest large language model accessible via API as of September 2026, delivering 1,107 tokens per second by vendor report and 440 tokens per second at median latency per OpenRouter's third-party telemetry. The model uses a diffusion-based architecture that processes token blocks in parallel rather than sequentially, giving it a structural speed advantage over autoregressive rivals like GPT-5.6 Luna and Claude Haiku 4.5. Mercury 2.5 is priced at $0.20 per million input tokens and $0.75 per million output tokens, undercutting competitors on output cost, though an 80% launch discount expired on 8 September 2026. GPT-5.6 Luna remains the preferred option for general-purpose chat due to its larger context window and broader ecosystem maturity, despite its slower speed of 129 tokens per second. Claude Haiku 4.5 and Gemini 3.5 Flash-Lite occupy the middle ground, with Flash-Lite offering competitive speed at a higher output price and Haiku favoured for Anthropic's instruction-following behaviour rather than throughput.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in