Inference engine compatibility
CUDA/ROCm support for the tools people actually use to run models locally or in production — vLLM, llama.cpp, Ollama, and ExLlama — as distinct from the training-focused PyTorch/TensorFlow tables elsewhere on this site. vLLM ships real pinned CUDA-wheel releases, so it gets a version history below like those two; the other three don't version that way (see each section's note), so they get a single current-state entry instead.
vLLM
| Version | Released | CUDA | ROCm | Status |
|---|---|---|---|---|
| 0.28.0 | 2026-08-26 | 13.0 | 7.0, 7.2.1 | current |
| 0.20.0 | 2026-04-27 | 13.0 | — | historical |
| 0.11.1 | 2025-11-18 | 12.9 | — | historical |
| 0.9.0 | 2025-05-15 | 12.8 | — | historical |
| 0.8.5 | 2025-04-28 | 12.4 | — | historical |
llama.cpp
rolling release (no numbered versions)current