Gemma 4 26B-A4B — VRAM requirements
Google · 25.81B params (3.8B active, mixture-of-experts) · up to 256K context
| Quantization | VRAM at 4K context | VRAM at 256K context |
|---|---|---|
| FP16 / BF16 | 60.2 GB | 110.3 GB |
| FP8 | 30.5 GB | 80.6 GB |
| INT8 | 30.5 GB | 80.6 GB |
| INT4 | 15.6 GB | 65.8 GB |
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.
Compatible GPUs at 256K context
FP16 / BF16 (110.3 GB)
FP8 (80.6 GB)
INT8 (80.6 GB)
INT4 (65.8 GB)
Sparse mixture-of-experts. **Both numbers in the model's name are rounded up from what Google's own spec table states**: that table gives 25.2B total and 3.8B active (8 active of 128 experts, plus 1 shared), not 26B/4B. paramsB here is the exact safetensors total (25.81B), which exceeds the card's 25.2B because it also includes the ~550M vision encoder; activeParamsB uses Google's stated 3.8B. The weights term uses the total, since all 128 experts must be VRAM-resident. maxContextLength 262,144 per the card. Pulled 2026-09-19.