Gemma 4 12B — VRAM requirements
Google · 11.96B params · up to 256K context
| Quantization | VRAM at 4K context | VRAM at 256K context |
|---|---|---|
| FP16 / BF16 | 29.2 GB | 138.6 GB |
| FP8 | 15.5 GB | 124.9 GB |
| INT8 | 15.5 GB | 124.9 GB |
| INT4 | 8.61 GB | 118.0 GB |
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.
Compatible GPUs at 256K context
FP16 / BF16 (138.6 GB)
FP8 (124.9 GB)
INT8 (124.9 GB)
INT4 (118.0 GB)
paramsB (11.96B) is the exact safetensors total from the Hugging Face API for google/gemma-4-12B-it (ungated); Google's card states 11.95B. Dense model. Architecture fields from config.json's text_config. maxContextLength 262,144 per Google's card ("the medium models support 256K"). Pulled 2026-09-19.