GLM 5.2

GLM 5.2 is Z.ai's MIT-licensed flagship, a sparse mixture-of-experts model with 753 billion total parameters and roughly 40 billion active per token, supporting a 1 million token input context and 128K output. Its IndexShare architecture reuses a single indexer across every four sparse attention layers, cutting per-token FLOPs by 2.9x at full context length, while speculative decoding via MTP, IndexShare, and KVShare adds a 20% acceptance-length gain. Built for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Features

On-demand Deployments

On-demand deployments let you run GLM 5.2 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

1M-Context Efficiency

IndexShare and KVShare cut per-token FLOPs by 2.9x at 1M context, keeping long-horizon agent runs and repository-scale engineering affordable.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image