Qwen3 8B — VRAM requirements
Alibaba · 8.19B params · up to 32K context
| Quantization | VRAM at 4K context | VRAM at 32K context |
|---|---|---|
| FP16 / BF16 | 19.5 GB | 24.4 GB |
| FP8 | 10.1 GB | 15.0 GB |
| INT8 | 10.1 GB | 15.0 GB |
| INT4 | 5.40 GB | 10.3 GB |
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.
Compatible GPUs at 32K context
FP16 / BF16 (24.4 GB)
FP8 (15.0 GB)
INT8 (15.0 GB)
INT4 (10.3 GB)
paramsB (8.19B) is the exact safetensors total from the Hugging Face API for Qwen/Qwen3-8B (ungated) — Qwen's own card rounds this to "8.2B". Architecture fields from that repo's config.json. maxContextLength is Qwen's documented native 32,768 (131,072 with YaRN), not config.json's 40,960. Pulled 2026-09-19.