DeepSeek V4 Pro
DeepSeek V4 Pro is the flagship of the MIT-licensed DeepSeek V4 family, a mixture-of-experts model with 1.6 trillion total parameters and roughly 49 billion active per token, served at a 164K context with a 384K output window. Its hybrid attention architecture combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) so that at long context it requires only 27% of the per-token inference FLOPs and 10% of the KV cache of DeepSeek V3.2 — making frontier-scale reasoning, coding, and agentic work practical at full context.
Features
On-demand Deployments
On-demand deployments let you run DeepSeek V4 Pro on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.
DocsCompressed Sparse Attention
At long context, CSA and HCA together need just 27% of the inference FLOPs and 10% of the KV cache of DeepSeek V3.2.
Docs