Chinese AI startup Moonshot AI has launched Kimi K3, a massive new open-weight large language model that the company says can rival the most advanced systems from US giants like OpenAI and Anthropic.[reference:0]
With 2.8 trillion parameters, Kimi K3 is the world's largest open-weight model, surpassing previous records held by DeepSeek's 1.6-trillion-parameter models.[reference:1][reference:2] The model is built on a Mixture of Experts (MoE) architecture featuring 896 experts, of which only 16 are activated per inference—a design that balances massive capacity with computational efficiency.[reference:3][reference:4] It also natively supports visual understanding and boasts a 1-million-token context window, allowing it to process the equivalent of a trilogy of novels in a single prompt.[reference:5][reference:6]
Frontier-Level Performance
Moonshot AI claims Kimi K3 delivers "frontier-level" performance across coding, reasoning, and knowledge work.[reference:7] While the company acknowledges the model still trails leading closed-source systems like Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol, it reportedly outperforms all other models in its evaluation suite.[reference:8]
Third-party benchmarks support these claims. Arena.ai ranked Kimi K3 first in the Frontend Code Arena with 1,679 points, surpassing Claude Fable 5—a 17-place jump from its predecessor Kimi K2.6.[reference:9] Artificial Analysis positioned the model at #3 on its Intelligence Index, with performance comparable to Opus 4.8 and GPT-5.5.[reference:10] In the Kernel Optimization Arena, K3's performance at maximum reasoning effort approached that of Fable 5, substantially outperforming Opus 4.8, GPT-5.6 Sol, and GPT-5.5.[reference:11]
Architectural Innovations
Kimi K3 incorporates two key architectural innovations. Kimi Delta Attention (KDA) is a hybrid linear attention mechanism that replaces the quadratic cost of traditional Transformer attention with linear scaling, reducing KV cache usage by 75% and boosting decoding throughput by up to 6x at 1-million-token context length.[reference:12] Attention Residuals (AttnRes) allows the model to selectively retrieve information across layers, preventing deeper layers from "forgetting" what was learned in earlier ones.[reference:13] Together, these innovations deliver a 2.5x improvement in training efficiency over the previous Kimi K2 generation.[reference:14]
Market Impact and Geopolitical Context
The launch comes at a sensitive moment. Just weeks earlier, the US government had temporarily forced Anthropic to withdraw its flagship Fable and Mythos models due to cybersecurity concerns.[reference:15] Kimi K3's rapid arrival suggests Chinese firms are successfully advancing independently despite US restrictions on hardware sales.[reference:16] The announcement triggered a sharp selloff in shares of Moonshot's domestic competitors Zhipu and MiniMax, which tumbled approximately 27% and 16% respectively in Hong Kong trading.[reference:17]
Kimi K3 is already available through Kimi.com, Kimi Work, Kimi Code, and the Kimi API.[reference:18] The model uses "max thinking effort" by default, with low- and high-effort modes planned for future updates.[reference:19] Full model weights are scheduled for public release by July 27, 2026, at which point researchers and developers worldwide will be able to download, run, and customize the system.[reference:20][reference:21]
Pricing is set to compete with Sonnet-class models: 3.00 per million (cache miss), and $15.00 per million output tokens.[reference:22]
What This Means for the AI Ecosystem
Kimi K3's open-weight nature is what truly sets it apart from US rivals. Unlike closed systems from OpenAI and Anthropic, anyone with sufficient hardware can run and modify Kimi K3.[reference:23] This model's arrival signals a new phase in the global AI race—one where open-source intelligence from China now competes directly with the most advanced closed-source systems from the West. As one researcher put it, the model represents a potential "Sputnik moment" for the industry.[reference:24]