vLLM 0.9.0

historical
Released
2025-05-15
CUDA
12.8
Python versions
>=3.9,<3.13
First release built with the default CUDA_VERSION bumped to 12.8.1 (from 12.4.1 in v0.8.5) per docker/Dockerfile's ARG CUDA_VERSION at this tag -- as with every entry in this collection, this is vLLM's own default build target, not an exhaustive compatibility list; the wheel auto-detects the installed CUDA rather than pinning to one version. Source: github.com/vllm-project/vllm at tag v0.9.0 (docker/Dockerfile, pyproject.toml) and the GitHub Releases API, verified 2026-09-06.

← All vLLM entries · All inference engines →