Alibaba Qwen3.8-Flash-Next and Google Gemini 3.7 Flash Reshape LLM Inference in 2026
Google launched Gemini 3.7 Flash on August 13, 2026, introducing configurable thinking budgets and hybrid reasoning into its fast inference model. Thirteen days later, Alibaba Cloud released Qwen3.8-Flash-Next, an open-weight model featuring 125 billion total parameters with only 6 billion activated per token. Alibaba's model pairs Gated DeltaNet recurrent linear attention with dynamic sparse attention to reduce compute costs at long context lengths, achieving over 175 tokens per second on standard datacenter hardware. Gemini 3.7 Flash, by contrast, retains full dense attention and relies on Google's TPU infrastructure and dynamic thinking budgets to manage memory efficiency, with a native context window of up to one million tokens. On pricing, Qwen3.8-Flash-Next is offered at $0.15 per million input tokens and $0.47 per million output tokens, while Gemini 3.7 Flash is positioned below $1.00 and $3.00 per million tokens for input and output respectively.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in