Kimi K3 Released: Moonshot AI's 2.8T Open-Weight Model Shakes the AI Frontier
Beijing, China — July 16, 2026 — Moonshot AI has officially unveiled Kimi K3, the company's most powerful flagship model to date and what it claims is the world's first open-source model in the 3-trillion-parameter class. The release marks a significant milestone in the global AI race, with an open-weight model now standing within striking distance of the best closed systems from OpenAI and Anthropic.
Launched across the Kimi platform, Kimi Code, and the API, K3 arrives with a 1-million-token context window, native multimodal understanding (text, image, and video), and a Mixture-of-Experts architecture that activates only a fraction of its 2.8 trillion parameters per token — making it both powerful and efficient.
Kimi K3 at a Glance: The Specs
| Developer | Moonshot AI (Beijing) |
|---|---|
| Release Date | July 16, 2026 |
| Total Parameters | ~2.8 trillion (MoE) |
| Active Experts | 16 of 896 per token |
| Context Window | 1,048,576 tokens (1M) |
| Modalities | Native text, image, and video |
| Variants | K3 Max (chat/agents), K3 Swarm Max (parallel processing) |
| API Pricing | $3/M input, $0.30/M cached, $15/M output |
| Open Weights | Promised by July 27, 2026 |
Frontier-Level Performance, Open-Weight Access
On the independent Artificial Analysis Intelligence Index v4.1, Kimi K3 scored 57.1, ranking as the #4 tested configuration and effectively the #3 model family globally — trailing only GPT-5.6 Sol Max (58.9) and Claude Fable 5 (59.9). This places it roughly on par with Claude Opus 4.8 and GPT-5.5, a remarkable achievement for an open-weight model.
Moonshot's own benchmarks show K3 achieving 88.3 on Terminal-Bench 2.1, 77.8 on Program Bench, and 42.0 on SWE Marathon. Independent testers also noted that K3 uses approximately 21% fewer output tokens than its predecessor K2.6, improving efficiency alongside capability.
According to Fortune, analysts were not expecting China to produce a model as powerful as Fable until early 2027. K3's arrival months ahead of schedule signals a dramatic compression in the open-vs-closed frontier gap.
Three Architectural Breakthroughs
K3 is not simply a scaled-up K2. Moonshot introduced three core architectural innovations:
- Kimi Delta Attention (KDA): A hybrid linear attention mechanism that enables up to 6.3x faster decoding in million-token contexts, solving the traditional slowdown associated with ultra-long sequences.
- Attention Residuals: A novel information-flow mechanism that allows later transformer blocks to draw directly from earlier representations, rather than relying on a strictly sequential residual stream. This delivers approximately 25% higher training efficiency.
- Stable LatentMoE: An advanced Mixture-of-Experts framework that increases sparsity — efficiently routing through 16 of 896 experts per token. Combined with training improvements, Moonshot claims 2.5x better overall scaling efficiency compared to K2.
What Makes K3 a Market Mover
The pricing strategy is deliberate. At $3/M input and $15/M output, K3 is roughly one-fifth the cost of comparable frontier models while delivering near-equivalent performance. Moonshot's own analysis shows K3 matching Fable 5 on software engineering tasks at approximately 35% of the price.
This pricing, combined with the promise of downloadable weights by July 27, positions K3 as a genuine alternative for enterprises that have been defaulting to closed APIs out of habit. As noted by BenchLM, "the distance between open models and the closed Western frontier is now measured in months, not years."
The UK AI Security Institute quantified this trend, finding the open-vs-closed gap has narrowed to 4–7 months, down from 6–10 months through most of 2025.
How to Access Kimi K3 Today
- Web & App: Available now at kimi.com and on iOS/Android
- Developers: Via the Kimi API and OpenRouter
- Coding: Integrated into Kimi Code for terminal and IDE workflows
- Self-hosting: Wait for open weights release (July 27, 2026) — expect to need 64+ accelerators for serious deployment
FAQ: Kimi K3 Answered
What is Kimi K3 and when was it released?
Kimi K3 is Moonshot AI's third-generation flagship large language model, released on July 16, 2026. It features approximately 2.8 trillion parameters in a Mixture-of-Experts (MoE) architecture with a 1-million-token context window and native multimodal capabilities including text, image, and video understanding.
How many parameters does Kimi K3 have?
Kimi K3 has approximately 2.8 trillion total parameters in a Mixture-of-Experts design with 896 experts, of which about 16 activate per token. This makes it the largest open-weight model ever announced, surpassing DeepSeek V4 Pro's 1.6 trillion parameters.
Is Kimi K3 open source?
Kimi K3 is described as an open-weight model, but the full weights have not been released yet. Moonshot AI has promised to release the complete model weights by July 27, 2026, under a permissive license. Until then, it is available as a hosted model through the Kimi app, Kimi Code, and API.
How does Kimi K3 compare to GPT-5.6 and Claude Fable 5?
On the independent Artificial Analysis Intelligence Index v4.1, Kimi K3 scored 57.1, placing it effectively as the #3 model family globally — close behind GPT-5.6 Sol Max (58.9) and Claude Fable 5 (59.9). It ranked #1 on the Frontend Code Arena, surpassing Claude Fable 5, and achieved competitive results on software engineering benchmarks like DeepSWE at approximately 35% of Fable 5's price.
What is the pricing for Kimi K3 API?
Kimi K3 API pricing is $3.00 per million input tokens (cache miss), $0.30 per million cached input tokens, and $15.00 per million output tokens. This represents roughly a 5x increase over Kimi K2.6 pricing, reflecting its frontier-level capabilities.
What are the key architectural innovations in Kimi K3?
Kimi K3 introduces three key architectural innovations: Kimi Delta Attention (KDA) — a hybrid linear attention mechanism for efficient long-sequence processing; Attention Residuals — allowing later layers to draw from earlier representations; and Stable LatentMoE — increasing sparsity with 16 of 896 experts activating per token. Moonshot claims these deliver roughly 2.5x better scaling efficiency than K2.
External Resources & Further Reading
Official Sources
Benchmarks & Analysis
News Coverage
The Bottom Line
Kimi K3 does not dethrone the absolute frontier — Claude Fable 5 and GPT-5.6 Sol still lead on broad benchmarks. But it does something arguably more important: it proves that an open-weight model can credibly join the frontier conversation. For developers, enterprises, and researchers, that changes the cost-and-control calculus entirely.
With weights dropping July 27, the next chapter in open AI is about to begin. Whether you're building agents, coding tools, or research pipelines, K3 deserves a spot in your evaluation queue.

Comments
Post a Comment