DeepSeek V4 Flash
DeepSeek V4 Flash is the efficiency-focused member of the MIT-licensed DeepSeek V4 family — 284 billion total parameters with roughly 13 billion active per token, sharing V4 Pro's hybrid Compressed Sparse Attention and Heavily Compressed Attention architecture and the same 164K served context and 384K output window. The smaller active footprint keeps long-context requests fast and inexpensive, making it well-suited to high-throughput agentic pipelines, document-scale analysis, and latency-sensitive production workloads.
Features
On-demand Deployments
On-demand deployments let you run DeepSeek V4 Flash on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.
DocsLow-Latency Long Context
Just 13B active parameters keep long-context requests fast and cheap, without giving up the V4 attention architecture.
Docs