RTX 4060 Ti 16GB

NVIDIAconsumer
Overview
ArchitectureAda Lovelace (2023)
Compute capability8.9
Memory
VRAM16 GB GDDR6
Memory bandwidth288 GB/s
Cache32 MB
Compute — vector
FP3222.1 TFLOPS
Cores & clocks
Shader cores4,352
Peak clock2,540 MHz
Board & system
TDP165 W
PCIePCIe Gen4
MSRP$499

Compatible CUDA Toolkit versions

11.812.012.112.212.312.412.512.612.812.913.013.113.213.3

PyTorch on the RTX 4060 Ti 16GB

Check against your driver and Python version →

The newest PyTorch, 2.14.0, runs on this card with a plain pip install torch.

PyTorchCUDA builds that run on this cardPlain pip install torch
2.14.012.6, 13.0, 13.2works
2.13.012.6, 12.9, 13.0, 13.2works
2.12.012.6, 13.0, 13.2works
2.11.012.6, 12.8, 12.9, 13.0works
2.10.012.6, 12.8, 12.9, 13.0works
2.9.012.6, 12.8, 12.9, 13.0works
2.8.012.6, 12.8, 12.9works
2.7.011.8, 12.6, 12.8works
2.6.011.8, 12.4, 12.6works
2.5.011.8, 12.1, 12.4works
2.4.011.8, 12.1, 12.4works
2.3.011.8, 12.1works
2.2.011.8, 12.1works
2.1.011.8, 12.1works
2.0.011.7, 11.8works
1.13.011.6, 11.7works
1.12.011.3, 11.6—
1.11.011.3—

From each build's compiled architecture list (Linux x86_64 wheels). Your driver must also support the CUDA version — see compatible CUDA versions.

Libraries on the RTX 4060 Ti 16GB

Whether each library's own requirements cover this card. Each row links to the version and source it was checked against on library support.

LibrarySupportedDetails
vLLM 0.30.0Yes
FlashAttention-2 2.8.3.post1Yes
FlashAttention-3 main (2026-09-27)Nono build for compute capability 8.9
bitsandbytes 0.50.2Yes
TensorRT 11.3.0Yes
llama.cpp master (2026-09-27)Yes

Native low-precision formats

Whether the card's matrix hardware runs these formats natively. Without native support a format can still work, but through slower emulation or conversion, not at the card's rated speed.

BF16FP8FP4
NativeNativeNo

Source: NVIDIA TensorRT 11.3.0 support matrix (hardware precision table).

Added 2026-09-27: the cheapest 16 GB NVIDIA card of its generation, a common local-LLM pick. Sources, all NVIDIA: 4,352 CUDA cores, 2.54 GHz boost, "22 TFLOPS" shader and "353 AI TOPS" from the RTX 4060 family product page (nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4060-4060ti/), whose Total Graphics Power row reads "165 or 160" for the 16 GB / 8 GB models, in the same order as its memory row. 288 GB/s from NVIDIA's VRAM deep-dive (nvidia.com/en-us/geforce/news/rtx-40-series-vram-video-memory-explained/), which describes the RTX 4060 Ti and its 32 MB L2 as "an Ada GPU with 288 GB/sec of peak memory bandwidth". The launch article says the 16 GB model has "additional graphics memory but otherwise identical specifications" and starts at $499. fp32TFLOPS 22.1 is cores × 2 × 2.54 GHz, matching NVIDIA's rounded 22. int8TOPSSparse holds the "AI TOPS" figure, which on Ada is the FP8/INT8 rate with sparsity (the same reading as the RTX 4090 entry's 1,321). The card's PCIe interface is x8 electrically.

← All GPUs · Compare GPUs →