vLLM 0.20.0

historical
Released
2026-04-27
CUDA
13.0
Python versions
>=3.10,<3.15
First release built against CUDA 13 (default CUDA_VERSION 13.0.2, up from 12.9.1 through v0.19.1) per docker/Dockerfile's ARG CUDA_VERSION -- vLLM's jump to the CUDA 13 major line. Source: github.com/vllm-project/vllm at tag v0.20.0 (docker/Dockerfile, pyproject.toml) and the GitHub Releases API, verified 2026-09-06.

← All vLLM entries · All inference engines →