Gemma 4 12B — VRAM requirements

Google · 11.96B params · up to 256K context

QuantizationVRAM at 4K contextVRAM at 256K context
FP16 / BF1629.2 GB138.6 GB
FP815.5 GB124.9 GB
INT815.5 GB124.9 GB
INT48.61 GB118.0 GB
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.

Compatible GPUs at 256K context

FP16 / BF16 (138.6 GB)

FP8 (124.9 GB)

INT8 (124.9 GB)

INT4 (118.0 GB)

paramsB (11.96B) is the exact safetensors total from the Hugging Face API for google/gemma-4-12B-it (ungated); Google's card states 11.95B. Dense model. Architecture fields from config.json's text_config. maxContextLength 262,144 per Google's card ("the medium models support 256K"). Pulled 2026-09-19.

← Interactive VRAM calculator · All GPUs