1,107 Tokens Per Second: The LLM That Doesn't Type

On September 8, 2026, Inception Labs announced Mercury 2.5 — which the company describes as the largest diffusion language model ever trained. The headline number: 1,107 tokens per second on widely available NVIDIA GPUs, at quality the company says matches the cost-optimized frontier tier (GPT-5.6 Luna Low, Gemini 3.5 Flash-Lite, Claude Haiku 4.5). It's a 40% intelligence jump over Mercury 2, and it's aimed at one thing the whole industry pretends doesn't exist: the speed ceiling built into how every mainstream LLM writes. Every LLM you've ever used writes like a typewriter. It picks token 1,
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in