GLM 5.2
GLM 5.2 is Z.ai's MIT-licensed flagship, a sparse mixture-of-experts model with 753 billion total parameters and roughly 40 billion active per token, supporting a 1 million token input context and 128K output. Its IndexShare architecture reuses a single indexer across every four sparse attention layers, cutting per-token FLOPs by 2.9x at full context length, while speculative decoding via MTP, IndexShare, and KVShare adds a 20% acceptance-length gain. Built for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.
Features
On-demand Deployments
On-demand deployments let you run GLM 5.2 on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.
Docs1M-Context Efficiency
IndexShare and KVShare cut per-token FLOPs by 2.9x at 1M context, keeping long-horizon agent runs and repository-scale engineering affordable.
Docs