On July 27, 2026, Moonshot AI made the full Kimi K3 model weights and technical report publicly available.
The 2.8 trillion parameter AI model marks the company's largest release, now accessible to developers and researchers worldwide.
Kimi K3 stands out with a mixture of experts (MoE) architecture, activating 104 billion parameters per inference from its 896 experts, selecting 16 experts per token.
It also supports a massive context window of over 1 million tokens, enabling native visual understanding and processing of extremely long inputs.
Moonshot shared the model weights on Hugging Face and posted the full technical report on their GitHub under the Kimi K3 License.
This rollout finalizes the initial launch earlier this month when the company promised additional technical details and full weights by late July.
The new model architecture reportedly improves compute efficiency by 2.5 times compared to the previous Kimi K2 model, thanks to innovations like Kimi Delta Attention, Attention Residuals, and a built-in quantization system implemented natively.
Developers can deploy Kimi K3 through popular frameworks such as Transformers, vLLM, and SGLang, though the model's size demands substantial hardware resources, with Moonshot advising high-performance supernode configurations.



