Qwen3.8
Qwen3.8 is the open-weights checkpoint behind Qwen3.8-Max and the largest model Alibaba's Qwen team has released — 2.4 trillion total parameters with 95 billion active per forward pass. A 92-layer stack alternates Gated DeltaNet linear attention with gated full attention, and each block ends in a mixture-of-experts layer where 10 routed experts plus 1 shared fire per token out of 512. It supports a 262,144-token context and is the first Qwen-Max-class model published with open weights. Served in two sizes — the 2.4T-A95B flagship and a 27B variant.
Features
On-demand Deployments
On-demand deployments let you run Qwen3.8 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.
DocsGated DeltaNet Attention
Alternating linear and gated full attention lets 95B active parameters serve a 2.4T model, with near-linear compute scaling across long contexts.
Docs