Technical profile
What developers should know
What Moonshot released
Moonshot announced Kimi K3 on July 14, 2026 and made it available through Kimi, Kimi Work, Kimi Code, and the Kimi API. The official API identifier is kimi-k3.[3] [1]
The Kimi product currently exposes Low, High, and Max thinking modes. Moonshot describes K3 as its most capable overall model; that is a vendor product claim, not an independent ranking across all models.[2]
Architecture and context
K3 combines Kimi Delta Attention and Attention Residuals with Stable LatentMoE. Moonshot reports that each token activates 16 of 896 experts and that the model uses a one-million-token context window.[1]
Moonshot describes K3 as natively multimodal, accepting visual information alongside text. The launch post also says quantization-aware training begins at supervised fine-tuning, using MXFP4 weights and MXFP8 activations.[1]
How to interpret the performance claim
Moonshot reports frontier-level results across its evaluation suite, but also says K3’s overall performance still trails the most powerful proprietary models it tested. Results in the launch post use different agent harnesses for some models, so benchmark comparisons need task-level inspection rather than a single “best model” label.[1]
The launch post identifies sensitivity to incomplete thinking history, excessive proactiveness on underspecified tasks, and a remaining user-experience gap versus the strongest proprietary systems as limitations. Production users should pin a compatible harness and set explicit action boundaries.[1]
Choose the artifact
Official checkpoints
K3 launched as a hosted model. Moonshot said the full weights would be released by July 27, 2026; as of this page’s July 18 verification, that date was still in the future.[1]
Moonshot recommends a compatible harness that preserves the complete thinking history and advises against switching into K3 midway through a session. For self-hosting, its launch guidance recommends supernode configurations with at least 64 accelerators; treat that as vendor guidance for the full model, not a universal minimum for future quantized variants.[1]