MiniMax H3

MiniMax H3 is an omni-modal generation model that reads text, images, video, and audio as one unified context and returns video with native stereo sound — up to 15 seconds at 2K resolution. Its H3-VAE tokenizer delivers a 4x gain in effective sequence length, letting a single model span text-to-video and first/last-frame-to-video without separate pipelines. Well-suited to creative production, product animation, and any workflow that needs synchronized picture and sound in one pass.

Features

On-demand Deployments

On-demand deployments let you run MiniMax H3 on dedicated GPUs with Sciforium's high-performance serving stack. Submit text-to-video (t2va) or first/last-frame-to-video (fl2va) requests at 768P or 2K and receive generated video with audio.

Docs

Native Stereo Audio

Video and synchronized stereo sound are generated together in a single pass, with no separate audio model or post-production step.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image