Qwen3.8

Qwen3.8 is the open-weights checkpoint behind Qwen3.8-Max and the largest model Alibaba's Qwen team has released — 2.4 trillion total parameters with 95 billion active per forward pass. A 92-layer stack alternates Gated DeltaNet linear attention with gated full attention, and each block ends in a mixture-of-experts layer where 10 routed experts plus 1 shared fire per token out of 512. It supports a 262,144-token context and is the first Qwen-Max-class model published with open weights. Served in two sizes — the 2.4T-A95B flagship and a 27B variant.

Features

On-demand Deployments

On-demand deployments let you run Qwen3.8 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

Gated DeltaNet Attention

Alternating linear and gated full attention lets 95B active parameters serve a 2.4T model, with near-linear compute scaling across long contexts.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image