Zhipu AI Releases GLM 5.3 Flash: Frontier-Level AI at a Fraction of Rival Costs
Chinese AI firm Zhipu AI (Z.ai) publicly revealed and open-sourced its GLM 5.3 Flash model on 26 August 2026, after it had spent a week atop OpenRouter and OpenCode leaderboards under the anonymous name ox-alpha. The model is a 320-billion-parameter mixture-of-experts architecture with roughly 18 billion active parameters per token, supporting text, image, and video inputs. Z.ai says it achieves a 3x reduction in attention compute and a 4.44x reduction in KV cache size compared to GLM 5.3, making large-context inference significantly cheaper. Independent benchmarker Artificial Analysis scores it at 57.5 on its intelligence index — nearly level with Claude Opus 4.8's 57.3 — while its promotional API pricing sits at roughly 1/70th of Opus 4.8's input cost. The weights are released under an MIT licence, and Z.ai confirmed the anonymous pre-launch testing was conducted on Chinese AI chips using a custom SGLang-based inference stack.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in