Gemma 4 31B

Gemma 4 31B is Google DeepMind's 30.7 billion parameter dense multimodal model — not a mixture-of-experts design — accepting text and image input and returning text. It offers a 262,144-token context with a 32,768-token maximum output, a configurable thinking mode, native function calling, and multilingual coverage across more than 140 languages, all under an Apache 2.0 license. Its dense architecture, including a roughly 550M-parameter vision encoder, makes it the most hardware-frugal model in the library while still competing with far larger frontier systems.

Features

On-demand Deployments

On-demand deployments let you run Gemma 4 31B on dedicated GPUs with Sciforium's high-performance serving stack, with high reliability and no rate limits.

Docs

Dense Multimodal Efficiency

31B dense parameters run on modest hardware while handling text and image input with configurable reasoning depth.

Docs
MiniMax  M2.5
Kimi K2.5
GLM 5
DeepSeek V3.2
gpt-oss-120b
gpt-oss-20b
Qwen3 Instruct
Qwen3 Thinking
Qwen3 Coder
Qwen3.5
Qwen3 VL Instruct
Qwen3 ASR
Qwen-Image
Qwen-Image-Edit
Flux2
Stable Diffusion 3.5
Hunyuan Image
Z-Image
Wan2.2-I2V
Wan2.2-T2V
Hunyuan Image
Z-Image