Kimi K3
Kimi K3 is a frontier mixture-of-experts model with 2.8 trillion total parameters, activating 16 of 896 experts — roughly 104 billion parameters — per token, making it the largest open-weight model released to date. It is built on two architectural innovations developed at Moonshot AI: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, a drop-in replacement for residual connections that delivers consistent scaling gains. With a 1 million token context, native visual understanding, and an always-on thinking mode, it is particularly strong at navigating large repositories, using tools, and iterating against images, logs, tests, and runtime feedback.
Features
On-demand Deployments
On-demand deployments let you run Kimi K3 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.
DocsKimi Delta Attention
Hybrid linear attention paired with Attention Residuals sustains a 1M-token context with consistent scaling gains over standard residual stacks.
Docs