DeepSeek V4 Pro

DeepSeek V4 Pro is the flagship of the MIT-licensed DeepSeek V4 family, a mixture-of-experts model with 1.6 trillion total parameters and roughly 49 billion active per token, served at a 164K context with a 384K output window. Its hybrid attention architecture combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) so that at long context it requires only 27% of the per-token inference FLOPs and 10% of the KV cache of DeepSeek V3.2 — making frontier-scale reasoning, coding, and agentic work practical at full context.

Features

On-demand Deployments

On-demand deployments let you run DeepSeek V4 Pro on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

Compressed Sparse Attention

At long context, CSA and HCA together need just 27% of the inference FLOPs and 10% of the KV cache of DeepSeek V3.2.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image