RTX 5060 Ti 16GB

NVIDIAconsumer
Overview
ArchitectureBlackwell (2025)
Compute capability12.0
Memory
VRAM16 GB GDDR7
Memory bandwidth448 GB/s
Compute — vector
FP3223.7 TFLOPS
Compute — matrix / tensor
FP4 (MXFP4 / NVFP4)379.5 TFLOPS
Cores & clocks
Shader cores4,608
Peak clock2,570 MHz
Board & system
TDP180 W
PCIePCIe Gen5

Compatible CUDA Toolkit versions

12.812.913.013.113.213.3

PyTorch on the RTX 5060 Ti 16GB

Check against your driver and Python version →

The newest PyTorch, 2.14.0, runs on this card with a plain pip install torch.

PyTorchCUDA builds that run on this cardPlain pip install torch
2.14.013.0, 13.2works
2.13.012.9, 13.0, 13.2works
2.12.013.0, 13.2works
2.11.012.8, 12.9, 13.0works
2.10.012.8, 12.9, 13.0works
2.9.012.8, 12.9, 13.0works
2.8.012.8, 12.9works
2.7.012.8doesn't (CUDA 12.6 build)
2.6.0nonedoesn't (CUDA 12.4 build)
2.5.0nonedoesn't (CUDA 12.4 build)
2.4.0nonedoesn't (CUDA 12.1 build)
2.3.0nonedoesn't (CUDA 12.1 build)
2.2.0nonedoesn't (CUDA 12.1 build)
2.1.0nonedoesn't (CUDA 12.1 build)
2.0.0nonedoesn't (CUDA 11.7 build)
1.13.0nonedoesn't (CUDA 11.7 build)
1.12.0none—
1.11.0none—

From each build's compiled architecture list (Linux x86_64 wheels). Your driver must also support the CUDA version — see compatible CUDA versions.

Libraries on the RTX 5060 Ti 16GB

Whether each library's own requirements cover this card. Each row links to the version and source it was checked against on library support.

LibrarySupportedDetails
vLLM 0.30.0Yes
FlashAttention-2 2.8.3.post1Yes
FlashAttention-3 main (2026-09-27)Nono build for compute capability 12.0
bitsandbytes 0.50.2Yes
TensorRT 11.3.0Yes
llama.cpp master (2026-09-27)Yes

Native low-precision formats

Whether the card's matrix hardware runs these formats natively. Without native support a format can still work, but through slower emulation or conversion, not at the card's rated speed.

BF16FP8FP4
NativeNativeNative

Source: NVIDIA TensorRT 11.3.0 support matrix (hardware precision table).

Added 2026-09-27: the cheapest 16 GB Blackwell card, with native FP4. Sources, all NVIDIA: 4,608 CUDA cores, 2.57 GHz boost, "759 AI TOPS", "16 GB / 8 GB GDDR7", 128-bit and 448 GB/sec from the RTX 50 Series table on NVIDIA's GeForce compare page (nvidia.com/en-us/geforce/graphics-cards/compare/), which also confirms this site's RTX 5090 bandwidth (1,792 GB/sec). 180 W Total Graphics Power from the RTX 5060 family product page. fp32TFLOPS 23.7 is cores × 2 × 2.57 GHz. fp4TFLOPS 379.5 is half the 759 "AI TOPS", which on Blackwell is the FP4 rate with sparsity (the same reading as the RTX 5090 entry's 3,352 / 1,676). No MSRP recorded: not found on an NVIDIA page.

← All GPUs · Compare GPUs →