meta: title: "Gemma 4 31B IT" slug: "gemma-4-31b-it" provider: "Google" description: "Google's unified multimodal Gemma 4 dense model (31B) with native text, image, and audio, plus thinking mode and tool-use protocol." date_added: 2026-04-02 date_updated: 2026-05-11 difficulty: intermediate tasks: - multimodal - text performance_headline: "Unified multimodal model with structured thinking, function calling, dynamic vision resolution" related_recipes: [] hardware: h100: verified mi300x: verified mi325x: verified mi355x: verified trillium: verified ironwood: verified model: model_id: "google/gemma-4-31B-it" min_vllm_version: "0.19.1" architecture: dense parameter_count: "31B" active_parameters: "31B" context_length: 262144 base_args: [] base_env: {} dependencies: - note: "Audio extras — only needed when serving the audio modality" command: 'uv pip install "vllm[audio]"' optional: true features: tool_calling: description: "Enable automatic tool choice with Gemma 4 parser and chat template" args: - "--enable-auto-tool-choice" - "--tool-call-parser" - "gemma4" - "--chat-template" - "examples/tool_chat_template_gemma4.jinja" reasoning: description: "Enable structured thinking/reasoning output" args: - "--reasoning-parser" - "gemma4" text_only: description: "Skip loading the vision encoder for text-only workloads — frees VRAM for KV cache. Mutually exclusive with encoder_parallel." args: - "--language-model-only" spec_decoding: description: "MTP speculative decoding for accelerated inference" default_mode: mtp modes: mtp: label: "MTP" description: "Built-in Multi-Token Prediction, 4 draft tokens." args: - "--speculative-config" - '{"model":"google/gemma-4-31B-it-assistant","num_speculative_tokens":4}' dflash: label: "DFlash" description: "DFlash speculative decoding with an externally trained draft model; the served checkpoint is unchanged." variants: - default args: - "--speculative-config" - '{"model": "RedHatAI/gemma-4-31B-it-speculator.dflash", "num_speculative_tokens": 7, "method": "dflash"}' dspark: label: "DSpark" description: "DSpark drafting with an externally trained draft model; the served checkpoint is unchanged." variants: - default args: - "--speculative-config" - '{"model": "RedHatAI/gemma-4-31B-it-speculator.dspark", "num_speculative_tokens": 7, "method": "dspark"}' eagle3: label: "Eagle3" description: "Eagle3 speculative decoding with externally trained draft model; the served checkpoint is unchanged." variants: - default args: - "--speculative-config" - '{"model": "RedHatAI/gemma-4-31B-it-speculator.eagle3", "num_speculative_tokens": 3, "method": "eagle3"}' opt_in_features: - text_only - spec_decoding variants: default: precision: bf16 vram_minimum_gb: 75 description: "Full BF16 — single 80 GB NVIDIA GPU or 1x MI300X/MI325X/MI350X/MI355X" fp8: model_id: "RedHatAI/gemma-4-31B-it-FP8-dynamic" precision: fp8 vram_minimum_gb: 38 description: "FP8 (E4M3) linear weights with dynamic per-token activation quantization (vision tower stays BF16) — Hopper or Blackwell" extra_args: - "--kv-cache-dtype" - "fp8" nvfp4: model_id: "nvidia/gemma-4-31B-it-NVFP4" precision: nvfp4 vram_minimum_gb: 19 description: "NVIDIA NVFP4 quantized weights for Blackwell GPUs" extra_args: - "--kv-cache-dtype" - "fp8" redhat_nvfp4: model_id: "RedHatAI/gemma-4-31B-it-NVFP4" precision: nvfp4 label: "NVFP4 (RedHat)" vram_minimum_gb: 19 description: "RedHatAI compressed-tensors NVFP4 quantization for Blackwell GPUs" extra_args: - "--kv-cache-dtype" - "fp8" w4a16: model_id: "google/gemma-4-31B-it-qat-w4a16-ct" precision: int4 vram_minimum_gb: 20 description: "W4A16 QAT (4-bit weights, 16-bit activations, group_size=32) via compressed-tensors — runs on any GPU" redhat_fp8: model_id: "RedHatAI/gemma-4-31B-it-FP8-block" precision: fp8 label: "FP8 Block (RedHat)" vram_minimum_gb: 38 description: "RedHatAI FP8 block quantization" extra_args: - "--kv-cache-dtype" - "fp8" compatible_strategies: - single_node_tp - multi_node_tp hardware_overrides: {} strategy_overrides: single_node_tp: tp: 1 guide: | ## Overview [Gemma 4](https://ai.google.dev/gemma/docs) is Google's most capable open model family, featuring a unified multimodal architecture that natively processes text, images, and audio. Gemma 4 models support structured thinking/reasoning, function calling with a custom tool-use protocol, and dynamic vision resolution — all available through vLLM's OpenAI-compatible API. ### Key Features - **Multimodal**: Text + images natively (video via custom frame-extraction pipeline). The smaller E2B and E4B models also support audio. - **MoE variant**: 128 fine-grained experts with top-8 routing and custom GELU-activated FFN (Gemma 4 26B-A4B). - **Dual Attention**: Alternating sliding-window (local) and global attention with different head dimensions. - **Thinking Mode**: Structured reasoning via `<|channel>thought\n...` delimiters. - **Function Calling**: Custom tool-call protocol with dedicated special tokens. - **Dynamic Vision Resolution**: Per-request configurable vision token budget (70, 140, 280, 560, 1120 tokens). ### Supported Variants Dense: - `google/gemma-4-E2B-it` (effective 2B) - `google/gemma-4-E4B-it` (effective 4B) - `google/gemma-4-31B-it` (31B) MoE: - `google/gemma-4-26B-A4B-it` (26B total / 4B active) TPU support is provided through [vLLM TPU](https://github.com/vllm-project/tpu-inference) with recipes for [Trillium](https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/trillium/vLLM/Gemma4) and [Ironwood](https://github.com/AI-Hypercomputer/tpu-recipes/blob/main/inference/ironwood/vLLM/Gemma4/). ## Prerequisites ### pip (NVIDIA CUDA) ```bash uv venv source .venv/bin/activate uv pip install -U vllm --pre \ --extra-index-url https://wheels.vllm.ai/nightly/cu129 \ --extra-index-url https://download.pytorch.org/whl/cu129 \ --index-strategy unsafe-best-match ``` ### pip (AMD ROCm: MI300X, MI325X, MI350X, MI355X) Requires Python 3.12, ROCm 7.2.1, glibc >= 2.35 (Ubuntu 22.04+). ```bash uv venv --python 3.12 source .venv/bin/activate uv pip install vllm --pre \ --extra-index-url https://wheels.vllm.ai/rocm/nightly/rocm721 --upgrade ``` ### Docker ```bash docker pull vllm/vllm-openai:gemma4-0505-cu129 # NVIDIA Hopper (H100/H200, CUDA 12.9) docker pull vllm/vllm-openai:gemma4-0505-cu130 # NVIDIA Blackwell (B200/B300, CUDA 13.0) docker pull vllm/vllm-openai-rocm:latest # AMD ``` TPU images are published separately by [vllm-project/tpu-inference](https://github.com/vllm-project/tpu-inference); see the Trillium / Ironwood tpu-recipes below for the pinned tag. ## Deployment Configurations ### Quick Start (Single GPU) ```bash vllm serve google/gemma-4-E4B-it \ --max-model-len # up to 131072 ``` ### 31B Dense on 2xA100/H100 (TP=2, BF16) ```bash vllm serve google/gemma-4-31B-it \ --tensor-parallel-size 2 \ --max-model-len 32768 \ --gpu-memory-utilization 0.90 ``` ### 26B MoE on 1xA100/H100 (BF16) ```bash vllm serve google/gemma-4-26B-A4B-it \ --max-model-len 32768 \ --gpu-memory-utilization 0.90 ``` ### Full-Featured Server Launch Enables text, image, audio, thinking, and tool calling: ```bash vllm serve google/gemma-4-31B-it \ --tensor-parallel-size 2 \ --max-model-len 16384 \ --gpu-memory-utilization 0.90 \ --enable-auto-tool-choice \ --reasoning-parser gemma4 \ --tool-call-parser gemma4 \ --chat-template examples/tool_chat_template_gemma4.jinja \ --limit-mm-per-prompt '{"image": 4, "audio": 1}' \ --async-scheduling \ --host 0.0.0.0 \ --port 8000 ``` ### Docker (NVIDIA) ```bash docker run -itd --name gemma4 \ --ipc=host --network host --shm-size 16G --gpus all \ -v ~/.cache/huggingface:/root/.cache/huggingface \ vllm/vllm-openai:gemma4-0505-cu129 \ --model google/gemma-4-31B-it \ --tensor-parallel-size 2 \ --max-model-len 32768 \ --gpu-memory-utilization 0.90 \ --host 0.0.0.0 --port 8000 ``` Swap `vllm/vllm-openai:gemma4-0505-cu129` for `vllm/vllm-openai:gemma4-0505-cu130` on Blackwell (B200/B300). ### Docker (AMD MI300X/MI325X/MI350X/MI355X) ```bash docker run -itd --name gemma4-rocm \ --ipc=host --network=host --privileged \ --cap-add=CAP_SYS_ADMIN --device=/dev/kfd --device=/dev/dri \ --group-add=video --cap-add=SYS_PTRACE \ --security-opt=seccomp=unconfined --shm-size 16G \ -v ~/.cache/huggingface:/root/.cache/huggingface \ vllm/vllm-openai-rocm:latest \ --model google/gemma-4-31B-it \ --host 0.0.0.0 --port 8000 ``` ### Docker (Cloud TPU — Trillium / Ironwood) TPU uses the separate `vllm/vllm-tpu` image (no pip wheel). Pull the tag specified by the upstream [Trillium](https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/trillium/vLLM/Gemma4) or [Ironwood](https://github.com/AI-Hypercomputer/tpu-recipes/blob/main/inference/ironwood/vLLM/Gemma4/) recipe, then run: ```bash docker run -itd --name gemma4-tpu \ --privileged --network host --shm-size 16G \ -v /dev/shm:/dev/shm -e HF_TOKEN=$HF_TOKEN \ vllm/vllm-tpu:latest \ --model google/gemma-4-31B-it \ --tensor-parallel-size 8 \ --max-model-len 16384 \ --disable_chunked_mm_input \ --host 0.0.0.0 --port 8000 ``` ## Client Usage ### Text Generation ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") response = client.chat.completions.create( model="google/gemma-4-31B-it", messages=[{"role": "user", "content": "Write a poem about the ocean."}], max_tokens=512, temperature=0.7, ) print(response.choices[0].message.content) ``` ### Image Understanding ```python response = client.chat.completions.create( model="google/gemma-4-31B-it", messages=[{"role": "user", "content": [ {"type": "image_url", "image_url": {"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/3/3a/Cat03.jpg/1200px-Cat03.jpg"}}, {"type": "text", "text": "Describe this image in detail."}, ]}], max_tokens=1024, ) ``` ### Dynamic Vision Resolution Supported values: 70, 140, 280 (default), 560, 1120 tokens/image. ```bash vllm serve google/gemma-4-31B-it \ --mm-processor-kwargs '{"max_soft_tokens": 560}' ``` ### Audio (E2B / E4B) Requires `uv pip install "vllm[audio]"`. ```bash vllm serve google/gemma-4-E2B-it \ --max-model-len 8192 \ --limit-mm-per-prompt '{"image": 4, "audio": 1}' ``` ### Thinking Mode ```bash vllm serve google/gemma-4-31B-it \ --max-model-len 16384 \ --enable-auto-tool-choice \ --reasoning-parser gemma4 \ --tool-call-parser gemma4 \ --chat-template examples/tool_chat_template_gemma4.jinja ``` Enable thinking per-request via `extra_body={"chat_template_kwargs": {"enable_thinking": True}}`, or default-on with `--default-chat-template-kwargs '{"enable_thinking": true}'`. ### Structured Outputs vLLM guided decoding constrains output to a JSON schema. Include semantic instructions in the system prompt — the model does not see schema descriptions. ## Configuration Tips - Set `--max-model-len` to match your workload. - `--gpu-memory-utilization 0.90–0.95` maximizes KV cache. - Image-only workloads: pass `--limit-mm-per-prompt.audio 0`. - Text-only workloads: pass `--limit-mm-per-prompt '{"image": 0, "audio": 0}'` to skip MM profiling. - `--async-scheduling` improves throughput. - FP8 KV cache (`--kv-cache-dtype fp8`) saves ~50% KV memory. ## Quantized Variants Two pre-quantized checkpoints are available: - [`RedHatAI/gemma-4-31B-it-FP8-dynamic`](https://huggingface.co/RedHatAI/gemma-4-31B-it-FP8-dynamic) — FP8 (E4M3) linear weights with dynamic per-token activation quantization (vision tower stays BF16); runs on Hopper and Blackwell. - [`nvidia/gemma-4-31B-it-NVFP4`](https://huggingface.co/nvidia/gemma-4-31B-it-NVFP4) — NVFP4 (4-bit) weights; requires Blackwell (B200/B300). Pick them from the **Variant** dropdown above, or pass the repo id directly to `vllm serve`. ## Throughput vs Latency | Goal | TP | `--max-num-seqs` | Notes | |------|----|------------------|-------| | Max throughput | 1-2 | 256-512 | Best tok/s per GPU | | Min latency | 4-8 | 8-16 | Best TTFT/TPOT | | Balanced | 2 | 128 | Mixed workloads | ## Speculative Decoding (MTP) Enable the **Spec Decoding** feature toggle (above) or add `--speculative-config` manually to use MTP drafting with the [assistant model](https://huggingface.co/google/gemma-4-31B-it-assistant). Recommended `num_speculative_tokens`: 4–8 for this model. See the [Gemma 4 usage guide](../../Google/Gemma4) for details and benchmarks. > **Note:** MTP speculative decoding for Gemma 4 is only available on the vLLM nightly build — it has not yet landed in a stable release. Install via the nightly wheel (`uv pip install -U vllm --pre --extra-index-url https://wheels.vllm.ai/nightly/cu129 …`) or use the `vllm/vllm-openai:gemma4-0505-cu129` / `vllm/vllm-openai:gemma4-0505-cu130` images above; the standard `:latest` stable tag does not include this feature. ## References - [Model card](https://huggingface.co/google/gemma-4-31B-it) - [FP8 variant](https://huggingface.co/RedHatAI/gemma-4-31B-it-FP8-dynamic) - [NVFP4 variant](https://huggingface.co/nvidia/gemma-4-31B-it-NVFP4) - [Gemma docs](https://ai.google.dev/gemma/docs) - [vLLM Gemma 4 tool-call template](https://github.com/vllm-project/vllm/blob/main/examples/tool_chat_template_gemma4.jinja) - [TPU recipes: Trillium](https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/trillium/vLLM/Gemma4) - [TPU recipes: Ironwood](https://github.com/AI-Hypercomputer/tpu-recipes/blob/main/inference/ironwood/vLLM/Gemma4/)