MiniMax M3
MiniMax M3 is a frontier mixture-of-experts model with 427 billion total parameters and 26 billion active parameters, trained multimodally from step zero so image and video are first-class inputs rather than bolted-on adapters. MiniMax Sparse Attention (MSA) carries a full 1 million token context at 9x prefill and over 15x decode speedup versus dense baselines, and the model reaches 59.0% on SWE-Bench Pro — making it well-suited to project-level software engineering, long-horizon agent runs, and computer-use workflows.
Features
On-demand Deployments
On-demand deployments let you run MiniMax M3 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.
Docs1M-Token Sparse Attention
MiniMax Sparse Attention delivers 9x prefill and over 15x decode speedup versus dense baselines, holding a full 1M-token context at practical cost.
Docs