Z.ai Releases GLM-5.3-Flash: 320B Open-Weight Model Active on Just 18B Parameters
Z.ai launched GLM-5.3-Flash on August 26, 2026, a 320-billion-parameter Mixture-of-Experts model that activates only 18 billion parameters per token, significantly reducing compute costs. The model supports a one-million-token context window and natively handles text, images, video, and files within a single request. It uses a hybrid sparse and linear attention system that Z.ai claims delivers 3x less attention computation and a 4.4x smaller KV cache compared to the full GLM-5.3. Released under the MIT license with weights available on Hugging Face, it was notably served on Chinese-made AI chips rather than Nvidia hardware. Prior to its official announcement, the model was tested anonymously under the codename ox-alpha on OpenRouter and OpenCode, where it became the most-used model of that week before its origin was revealed.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in