vLLM 0.28.0

current
Released
2026-08-26
CUDA
13.0
ROCm
7.07.2.1
Python versions
>=3.10,<3.15
Latest stable release as of 2026-09-06. Default CUDA_VERSION is 13.0.3 (docker/Dockerfile's ARG CUDA_VERSION; a patch bump from 13.0.2, unchanged since v0.20.0). On AMD, vLLM's own docs state it "supports AMD GPUs with ROCm 6.3 or above," with pre-built ROCm wheels currently published for ROCm 7.0 and 7.2.1 -- rocmVersions here lists those two published-wheel versions specifically, not the full "6.3 or above" range the docs also describe as generally working. Source: github.com/vllm-project/vllm at tag v0.28.0 (docker/Dockerfile, pyproject.toml), docs.vllm.ai/en/stable/getting_started/installation/gpu/ for the ROCm support statement, and the GitHub Releases API, verified 2026-09-06.

← All vLLM entries · All inference engines →