MiniMax M3

MiniMax M3 is a frontier mixture-of-experts model with 427 billion total parameters and 26 billion active parameters, trained multimodally from step zero so image and video are first-class inputs rather than bolted-on adapters. MiniMax Sparse Attention (MSA) carries a full 1 million token context at 9x prefill and over 15x decode speedup versus dense baselines, and the model reaches 59.0% on SWE-Bench Pro — making it well-suited to project-level software engineering, long-horizon agent runs, and computer-use workflows.

Features

On-demand Deployments

On-demand deployments let you run MiniMax M3 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

1M-Token Sparse Attention

MiniMax Sparse Attention delivers 9x prefill and over 15x decode speedup versus dense baselines, holding a full 1M-token context at practical cost.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image