Qwen3 8B — VRAM requirements

Alibaba · 8.19B params · up to 32K context

QuantizationVRAM at 4K contextVRAM at 32K context
FP16 / BF1619.5 GB24.4 GB
FP810.1 GB15.0 GB
INT810.1 GB15.0 GB
INT45.40 GB10.3 GB
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.

Compatible GPUs at 32K context

FP16 / BF16 (24.4 GB)

FP8 (15.0 GB)

INT8 (15.0 GB)

INT4 (10.3 GB)

GTX 1080 Ti (11 GB, NVIDIA)Radeon RX 6700 XT (12 GB, AMD)Radeon RX 6750 XT (12 GB, AMD)Radeon RX 7700 XT (12 GB, AMD)Arc B580 (12 GB, Intel)TITAN V (12 GB, NVIDIA)TITAN X (Pascal) (12 GB, NVIDIA)Radeon Instinct MI25 (16 GB, AMD)Radeon Instinct MI6 (16 GB, AMD)Radeon RX 6800 (16 GB, AMD)Radeon RX 6800 XT (16 GB, AMD)Radeon RX 6900 XT (16 GB, AMD)Radeon RX 6950 XT (16 GB, AMD)Radeon RX 7600 XT (16 GB, AMD)Radeon RX 7800 XT (16 GB, AMD)Radeon RX 7900 GRE (16 GB, AMD)Radeon RX 9060 XT (16GB) (16 GB, AMD)Radeon RX 9070 (16 GB, AMD)Radeon RX 9070 XT (16 GB, AMD)Radeon VII (16 GB, AMD)Arc A770 (16GB) (16 GB, Intel)P100 SXM2 16GB (16 GB, NVIDIA)T4 (16 GB, NVIDIA)Radeon RX 7900 XT (20 GB, AMD)Radeon RX 7900 XTX (24 GB, AMD)K80 (24 GB, NVIDIA)M40 24GB (24 GB, NVIDIA)RTX 3090 (24 GB, NVIDIA)RTX 4090 (24 GB, NVIDIA)TITAN RTX (24 GB, NVIDIA)Instinct MI100 (32 GB, AMD)Radeon AI PRO R9700 (32 GB, AMD)Radeon Instinct MI50 (32 GB, AMD)RTX 5090 (32 GB, NVIDIA)V100 SXM2 32GB (32 GB, NVIDIA)L40S (48 GB, NVIDIA)RTX 6000 Ada Generation (48 GB, NVIDIA)Instinct MI210 (64 GB, AMD)A100 SXM4 80GB (80 GB, NVIDIA)H100 SXM5 80GB (80 GB, NVIDIA)RTX PRO 6000 Blackwell (96 GB, NVIDIA)Instinct MI250X (128 GB, AMD)Instinct MI300A (128 GB, AMD)Data Center GPU Max 1550 (128 GB, Intel)H200 SXM 141GB (141 GB, NVIDIA)B200 (180 GB, NVIDIA)Instinct MI300X (192 GB, AMD)Instinct MI325X (256 GB, AMD)Instinct MI350X (288 GB, AMD)Instinct MI355X (288 GB, AMD)B300 (Blackwell Ultra) (288 GB, NVIDIA)Instinct MI455X (432 GB, AMD)
paramsB (8.19B) is the exact safetensors total from the Hugging Face API for Qwen/Qwen3-8B (ungated) — Qwen's own card rounds this to "8.2B". Architecture fields from that repo's config.json. maxContextLength is Qwen's documented native 32,768 (131,072 with YaRN), not config.json's 40,960. Pulled 2026-09-19.

← Interactive VRAM calculator · All GPUs