DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-focused member of the MIT-licensed DeepSeek V4 family — 284 billion total parameters with roughly 13 billion active per token, sharing V4 Pro's hybrid Compressed Sparse Attention and Heavily Compressed Attention architecture and the same 164K served context and 384K output window. The smaller active footprint keeps long-context requests fast and inexpensive, making it well-suited to high-throughput agentic pipelines, document-scale analysis, and latency-sensitive production workloads.

Features

On-demand Deployments

On-demand deployments let you run DeepSeek V4 Flash on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

Low-Latency Long Context

Just 13B active parameters keep long-context requests fast and cheap, without giving up the V4 attention architecture.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image