Gemma 4 26B-A4B — VRAM requirements

Google · 25.81B params (3.8B active, mixture-of-experts) · up to 256K context

QuantizationVRAM at 4K contextVRAM at 256K context
FP16 / BF1660.2 GB110.3 GB
FP830.5 GB80.6 GB
INT830.5 GB80.6 GB
INT415.6 GB65.8 GB
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.

Compatible GPUs at 256K context

FP16 / BF16 (110.3 GB)

FP8 (80.6 GB)

INT8 (80.6 GB)

INT4 (65.8 GB)

Sparse mixture-of-experts. **Both numbers in the model's name are rounded up from what Google's own spec table states**: that table gives 25.2B total and 3.8B active (8 active of 128 experts, plus 1 shared), not 26B/4B. paramsB here is the exact safetensors total (25.81B), which exceeds the card's 25.2B because it also includes the ~550M vision encoder; activeParamsB uses Google's stated 3.8B. The weights term uses the total, since all 128 experts must be VRAM-resident. maxContextLength 262,144 per the card. Pulled 2026-09-19.

← Interactive VRAM calculator · All GPUs