Best GPU for local LLMs

A curated view of this site's own GPU and VRAM calculator data, not a new dataset. VRAM is the constraint that actually matters for local LLM inference — not raw TFLOPS — because a model that doesn't fit in VRAM simply won't load. Memory bandwidth is the second-order factor: it's usually the real ceiling on tokens/second once a model fits.

By VRAM budget

"Fits" below means the single largest current-generation model in this site's VRAM calculator collection that comes in under each card's VRAM at INT4 quantization and 4K context — the same formula the calculator itself uses, not a separate rule of thumb. Check a specific model + context length combination on that page directly; this table is a starting point, not a substitute.

GPUVendorVRAMMSRPLargest fit (INT4, 4K context)
Arc A380Intel6 GB$139Qwen3 8B (8.19B)
Radeon RX 6600AMD8 GB—Ornith 1.5 9B (9.65B)
Radeon RX 6650 XTAMD8 GB—Ornith 1.5 9B (9.65B)
Radeon RX 9060 XT (8GB)AMD8 GB—Ornith 1.5 9B (9.65B)
Arc A750Intel8 GB$289Ornith 1.5 9B (9.65B)
GTX 1080 TiNVIDIA11 GB$699Gemma 4 12B (11.96B)
Radeon RX 6700 XTAMD12 GB—Gemma 4 12B (11.96B)
Radeon RX 6750 XTAMD12 GB—Gemma 4 12B (11.96B)
Radeon RX 7700 XTAMD12 GB—Gemma 4 12B (11.96B)
Arc B580Intel12 GB$249Gemma 4 12B (11.96B)
TITAN VNVIDIA12 GB$2,999Gemma 4 12B (11.96B)
TITAN X (Pascal)NVIDIA12 GB$1,200Gemma 4 12B (11.96B)
Radeon RX 6800AMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 6800 XTAMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 6900 XTAMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 6950 XTAMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 7600 XTAMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 7800 XTAMD16 GB$499Gemma 4 26B-A4B (25.81B)
Radeon RX 7900 GREAMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 9060 XT (16GB)AMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 9070AMD16 GB—Gemma 4 26B-A4B (25.81B)
Radeon RX 9070 XTAMD16 GB$599Gemma 4 26B-A4B (25.81B)
Radeon VIIAMD16 GB—Gemma 4 26B-A4B (25.81B)
Arc A770 (16GB)Intel16 GB$349Gemma 4 26B-A4B (25.81B)
Radeon RX 7900 XTAMD20 GB—Qwen3 32B (32.76B)
Radeon RX 7900 XTXAMD24 GB$999Ornith 1.5 35B-A3B (35.95B)
RTX 3090NVIDIA24 GB$1,499Ornith 1.5 35B-A3B (35.95B)
RTX 4090NVIDIA24 GB$1,599Ornith 1.5 35B-A3B (35.95B)
TITAN RTXNVIDIA24 GB$2,499Ornith 1.5 35B-A3B (35.95B)
Radeon AI PRO R9700AMD32 GB—Kimi Linear 48B-A3B (49.12B)
RTX 5090NVIDIA32 GB$1,999Kimi Linear 48B-A3B (49.12B)
RTX 6000 Ada GenerationNVIDIA48 GB$6,800Kimi Linear 48B-A3B (49.12B)
RTX PRO 6000 BlackwellNVIDIA96 GB—Mistral Small 4 119B (119.4B)

The 48GB and 96GB workstation cards show the same "largest fit" above — that's a gap in this site's tracked model list, not a data error: the next model up from Qwen2.5 72B is Llama 3.1 405B, which needs well over 96GB even at INT4, so nothing in between currently separates those two cards here.

Picks by use case

Best value entry point: Radeon RX 9070 XT or Intel Arc B580

Under $600, both give real headroom over the 8GB cards at the bottom of the table above — the RX 9070 XT (16GB, $599) is the newer of the two and got unusually fast official ROCm support for a just-launched card; the Arc B580 (12GB, $249) is the cheapest way into double-digit VRAM at all, though Intel's software stack for local LLM tools is less mature than CUDA or ROCm — expect more setup friction.

The sweet spot: RTX 3090 (24GB)

RTX 3090's 24GB and full CUDA support (plus the biggest secondhand market of any card on this page, since NVIDIA no longer sells it new) make it the card most local-LLM hobbyists actually land on — MSRP shown above is the original 2020 launch price, not current secondhand pricing, which this site doesn't track. The RX 7900 XTX matches its 24GB at a similar or lower price if you're committed to ROCm; see ROCm on consumer GPUs for what that trade-off actually involves. The RTX 4090 has more compute and bandwidth than either at the same 24GB, if budget allows the premium.

No compromises: RTX PRO 6000 Blackwell (96GB) or RTX 6000 Ada (48GB)

Both are workstation cards, not consumer GPUs, and priced accordingly (unofficial street pricing for the RTX PRO 6000 Blackwell has swung well above its original launch price — see that page's notes). What they buy: enough VRAM to run this site's largest tracked models directly rather than working around VRAM limits with aggressive quantization or a multi-GPU split (which this table doesn't model at all — every row above assumes a single card).

← VRAM calculator · Compare two GPUs directly →