Z.ai Confirms Mystery Model Ox Alpha Was Secretly Tested GLM-5.3-Flash
Z.ai confirmed on August 26 that a mysterious model called Ox Alpha, which appeared anonymously on OpenCode and OpenRouter on August 20, was actually its new GLM-5.3-Flash model being tested in stealth to gather real-world developer feedback. The model is a 320-billion-parameter Mixture-of-Experts architecture with only 18 billion parameters active per token and a 1-million-token context window. GLM-5.3-Flash is the first model in the GLM-5 series to natively support text, image, and video input at the architecture level, and is released as open weights under the MIT license. It is also claimed by Z.ai to be the first open-source frontier model combining sparse and linear attention in a single architecture, reducing attention compute by 3x and KV cache size by 4.4x compared to its predecessor. Z.ai states the model delivers performance comparable to more expensive frontier models at roughly one-tenth the cost of its predecessor, though benchmark figures are vendor-reported and have not been independently verified.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in