Moonshot AI Releases Kimi K3, a 2.8T-Parameter Open-Weights MoE Model
Moonshot AI has launched Kimi K3, an open-weights large language model featuring 2.8 trillion total parameters built on a Mixture-of-Experts architecture that activates 104 billion parameters per token. The model delivers a reported 2.5x improvement in scaling efficiency over its predecessor, Kimi K2, achieved through architectural refinements rather than simply increased compute. Key innovations include Kimi Delta Attention for better information flow across deep layers and a sparse routing system that activates 16 out of 896 experts per token. Kimi K3 supports a 1-million-token context window and was trained using agentic reinforcement learning with sandbox states, making it suited for complex multi-step coding and long-context reasoning tasks. By releasing the model weights publicly, Moonshot AI enables developers and researchers to fine-tune, inspect, and deploy K3 without relying on proprietary API access.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in