Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Preview of Qwen4 Architecture
Alibaba released Qwen3.8-Flash-Next on August 26, 2026, as an open-weight public preview of the architecture planned for its upcoming Qwen4 model family. The sparse mixture-of-experts model carries 125 billion total parameters but activates only around 6 billion per token, plus a 51-billion-parameter N-gram embedding table used for lookups rather than computation. It supports a native context window of 262,144 tokens, extendable to one million, and handles text, image, and video inputs. Alibaba claims the model was trained at roughly one-ninth the cost of its predecessor Qwen3.7-Plus, using a hybrid Gated DeltaNet and Qwen Sparse Attention architecture alongside the Muon optimizer. While Flash-Next is better suited for large-document agents and cost-sensitive API workloads, its memory footprint makes it impractical for single consumer GPU deployment, where the denser Qwen3.8-27B remains the more practical option.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in