Kimi K3

Kimi K3 is a frontier mixture-of-experts model with 2.8 trillion total parameters, activating 16 of 896 experts — roughly 104 billion parameters — per token, making it the largest open-weight model released to date. It is built on two architectural innovations developed at Moonshot AI: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, a drop-in replacement for residual connections that delivers consistent scaling gains. With a 1 million token context, native visual understanding, and an always-on thinking mode, it is particularly strong at navigating large repositories, using tools, and iterating against images, logs, tests, and runtime feedback.

Features

On-demand Deployments

On-demand deployments let you run Kimi K3 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

Kimi Delta Attention

Hybrid linear attention paired with Attention Residuals sustains a 1M-token context with consistent scaling gains over standard residual stacks.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image