Library support by GPU

Every card on this site against the GPU requirements of the libraries people install next to PyTorch. Each rule comes from the library's own documentation or build configuration at the version shown, linked below.

✓ supported · partly = some features or backends only · ✗ not supported · ? not enough data. For which PyTorch builds run on a card, see its GPU page.

GPUvLLMFlashAttention-2FlashAttention-3bitsandbytesTensorRTllama.cpp
B300 (Blackwell Ultra) NVIDIA · 10.0✓✓✗✓✓✓
RTX 5090 NVIDIA · 12.0✓✓✗✓✓✓
RTX PRO 6000 Blackwell NVIDIA · 12.0✓✓✗✓✓✓
B200 NVIDIA · 10.0✓✓✗✓✓✓
H200 SXM 141GB NVIDIA · 9.0✓✓✓✓✓✓
L40S NVIDIA · 8.9✓✓✗✓✓✓
H100 SXM5 80GB NVIDIA · 9.0✓✓✓✓✓✓
RTX 4090 NVIDIA · 8.9✓✓✗✓✓✓
RTX 6000 Ada Generation NVIDIA · 8.9✓✓✗✓✓✓
A100 SXM4 80GB NVIDIA · 8.0✓✓✗✓✓✓
RTX 3090 NVIDIA · 8.6✓✓✗✓✓✓
T4 NVIDIA · 7.5✓✗✗✓✓✓
TITAN RTX NVIDIA · 7.5✓✗✗✓✓✓
GTX 1080 Ti NVIDIA · 6.1✗✗✗partly✗✓
TITAN V NVIDIA · 7.0✗✗✗partly✗✓
V100 SXM2 32GB NVIDIA · 7.0✗✗✗partly✗✓
P100 SXM2 16GB NVIDIA · 6.0✗✗✗partly✗✓
Tesla P40 NVIDIA · 6.1✗✗✗partly✗✓
TITAN X (Pascal) NVIDIA · 6.1✗✗✗partly✗✓
M40 24GB NVIDIA · 5.2✗✗✗✗✗✓
K80 NVIDIA · 3.7✗✗✗✗✗partly
Instinct MI455X AMD??????
Instinct MI350X AMD · gfx950✓✓✗✓✗✓
Instinct MI355X AMD · gfx950✓✓✗✓✗✓
Radeon AI PRO R9700 AMD · gfx1201✓✗✗✓✗✓
Radeon RX 9060 XT (16GB) AMD · gfx1200✓✗✗✓✗✓
Radeon RX 9060 XT (8GB) AMD · gfx1200✓✗✗✓✗✓
Radeon RX 9070 AMD · gfx1201✓✗✗✓✗✓
Radeon RX 9070 XT AMD · gfx1201✓✗✗✓✗✓
Instinct MI325X AMD · gfx942✓✓✗✓✗✓
Radeon RX 7600 XT AMD · gfx1102✗✗✗✓✗✓
Instinct MI300A AMD · gfx942✓✓✗✓✗✓
Instinct MI300X AMD · gfx942✓✓✗✓✗✓
Radeon RX 7700 XT AMD · gfx1101✓✗✗✓✗✓
Radeon RX 7800 XT AMD · gfx1101✓✗✗✓✗✓
Radeon RX 7900 GRE AMD · gfx1100✓✗✗✓✗✓
Instinct MI210 AMD · gfx90a✓✓✗✓✗✓
Radeon RX 6650 XT AMD · gfx1032✗✗✗✓✗partly
Radeon RX 6750 XT AMD · gfx1031✗✗✗✓✗partly
Radeon RX 6950 XT AMD · gfx1030✗✗✗✓✗✓
Radeon RX 7900 XT AMD · gfx1100✓✗✗✓✗✓
Radeon RX 7900 XTX AMD · gfx1100✓✗✗✓✗✓
Instinct MI250X AMD · gfx90a✓✓✗✓✗✓
Radeon RX 6600 AMD · gfx1032✗✗✗✓✗partly
Radeon RX 6700 XT AMD · gfx1031✗✗✗✓✗partly
Instinct MI100 AMD · gfx908✗✗✗✓✗✓
Radeon RX 6800 AMD · gfx1030✗✗✗✓✗✓
Radeon RX 6800 XT AMD · gfx1030✗✗✗✓✗✓
Radeon RX 6900 XT AMD · gfx1030✗✗✗✓✗✓
Radeon VII AMD · gfx906✗✗✗✗✗partly
Radeon Instinct MI50 AMD · gfx906✗✗✗✗✗partly
Radeon Instinct MI60 AMD · gfx906✗✗✗✗✗partly
Radeon Instinct MI25 AMD · gfx900✗✗✗✗✗partly
Radeon Instinct MI6 AMD · gfx803✗✗✗✗✗partly
Radeon Instinct MI8 AMD · gfx803✗✗✗✗✗partly
Arc B580 Intel✓✗✗✓✗✓
Data Center GPU Max 1550 Intel✓✗✗✓✗✓
Arc A380 Intel✓✗✗✓✗✓
Arc A750 Intel✓✗✗✓✗✓
Arc A770 (16GB) Intel✓✗✗✓✗✓

vLLM 0.30.0

Rule: NVIDIA compute capability 7.5 or higher; AMD gfx90a, gfx942, gfx950, gfx1100, gfx1101, gfx1150, gfx1151, gfx1200, gfx1201; Intel supported.

AMD: ROCm 6.3 or newer (MI350 needs ROCm 7.0+, Ryzen AI MAX / AI 300 need 7.0.2+). Pre-built wheels for ROCm 7.0 and 7.2.1.

Intel: XPU backend: Intel Data Center GPU and Arc GPUs.

Requirements as stated in vLLM's own installation docs at the v0.30.0 tag. The NVIDIA minimum has been compute capability 7.5 since at least v0.20.0 (checked at v0.20.0, 0.24.0, 0.26.0-0.29.0), so the V100 (7.0) is not supported. The AMD list is the docs' "GPU:" requirement line: MI200s (gfx90a), MI300 (gfx942), MI350 (gfx950), Radeon RX 7900 series (gfx1100/1101), Radeon RX 9000 series (gfx1200/1201), Ryzen AI MAX / AI 300 (gfx1151/1150).

Checked 2026-09-27 against: gpu.cuda.inc.md, gpu.rocm.inc.md, gpu.xpu.inc.md

FlashAttention-2 2.8.3.post1

Rule: NVIDIA builds compiled for compute capability 8.0, 9.0, 10.0, 12.0 (and newer minor versions of each); AMD gfx90a, gfx942, gfx950; no Intel support.

AMD: Composable Kernel backend, ROCm 6.0+. RDNA 3/4 and Strix targets are allowed on the main branch but not in a release yet.

compiledFor is the default FLASH_ATTN_CUDA_ARCHS in setup.py at the v2.8.3.post1 tag (sm_100/sm_120 need a CUDA 12.8+ build). The README text names Ampere, Ada and Hopper; the build also covers Blackwell. Ada (8.9) runs the sm_80 code under CUDA's same-major compatibility. Turing (T4, RTX 20xx) is not supported here; the README points to the separate flash-attention-turing project. The AMD list is setup.py's allowed ROCm archs at the same tag. The main branch (2026-09-27) also allows gfx1100-1102, gfx1150/1151 and gfx1200/1201, which the README already advertises, but no release includes them yet.

Checked 2026-09-27 against: setup.py, README.md

FlashAttention-3 main (2026-09-27)

Rule: NVIDIA builds compiled for compute capability 9.0 (and newer minor versions of each); no AMD support; no Intel support.

The README's stated requirement is "H100 / H800 GPU, CUDA >= 12.3". FlashAttention-3 is built from the repo's hopper/ directory and is Hopper-only. FlashAttention-4 (the flash-attn-4 package, 4.0.0b32 on 2026-09-27) is a beta and is not tracked here yet.

Checked 2026-09-27 against: README.md

bitsandbytes 0.50.2

Rule: NVIDIA builds compiled for compute capability 6.0, 7.0, 7.5, 8.0, 8.6, 8.9, 9.0, 10.0, 12.0 (and newer minor versions of each); AMD gfx908, gfx90a, gfx942, gfx950, gfx1250, gfx1010, gfx1011, gfx1012, gfx1030, gfx1031, gfx1032, gfx1033, gfx1034, gfx1035, gfx1036, gfx1100, gfx1101, gfx1102, gfx1103, gfx1150, gfx1151, gfx1152, gfx1153, gfx1200, gfx1201; Intel supported.

LLM.int8() needs compute capability 7.5 or higher.

AMD: Pre-built for ROCm 7.14.0 and 10.0.0 (Linux x86_64); older ROCm builds cover fewer RDNA targets.

Intel: Intel XPU through PyTorch's XPU backend (PyTorch 2.6+).

From the installation docs (2026-09-27): NVIDIA compute capability 6.0+, with LLM.int8() needing 7.5+ (8-bit optimizers and NF4/FP4 quantization need 6.0+). compiledFor is the union of the Linux x86_64 PyPI wheel targets across its CUDA builds (11.8-12.6: sm60-sm90; 12.8: sm70-sm120; 13.0-13.2: sm75-sm120). The AMD list is the ROCm 7.14.0 / 10.0.0 Linux wheel targets. The RX 6700 XT's gfx1031 and RX 6600's gfx1032 are included there, though ROCm itself doesn't officially support those cards.

Checked 2026-09-27 against: installation.mdx

TensorRT 11.3.0

Rule: NVIDIA compute capability 7.5 or higher; no AMD support; no Intel support.

NVIDIA's TensorRT 11.3.0 support matrix: "TensorRT supports NVIDIA hardware with compute capability SM 7.5 or higher." NVIDIA-only.

Checked 2026-09-27 against: support-matrix.html

llama.cpp master (2026-09-27)

Rule: NVIDIA compute capability 5.0 or higher; AMD GPUs current ROCm supports (others: Vulkan only); Intel supported.

NVIDIA: CUDA backend. Compute capability 5.0-7.0 needs a CUDA 12 build (CUDA 13 builds start at 7.5). Older cards can use the Vulkan backend.

AMD: HIP backend on cards ROCm supports; the Vulkan backend works on the rest.

Intel: SYCL or Vulkan backend.

llama.cpp has several GPU backends, so it runs on almost any card; what varies is which backend. The CUDA backend's default build targets are 50/61/70-virtual (CUDA < 13 only), 75/80-virtual, 86-real, 89-real, 90-virtual and 120a-real (CUDA 12.8+), per ggml-cuda's CMakeLists.txt.

Checked 2026-09-27 against: CMakeLists.txt, build.md