RTX 5090

NVIDIAconsumer
Overview
ArchitectureBlackwell (2025)
Process nodeTSMC 4N
Transistor count92.2B
Compute capability12.0
Memory
VRAM32 GB GDDR7
Memory bandwidth1,792 GB/s
Cache96 MB
Compute — vector
FP32104.8 TFLOPS
Compute — matrix / tensor
TF32104.8 TFLOPS
BF16209.5 TFLOPS
FP16 (dense / sparse)209.5 / 419 TFLOPS
FP8 (dense / sparse)419 / 838 TFLOPS
FP4 (MXFP4 / NVFP4)1,676 TFLOPS
INT8 (dense / sparse)838 / 1,676 TOPS
Cores & clocks
Streaming Multiprocessors170
Shader cores21,760
Matrix / Tensor cores680
Peak clock2,407 MHz
Board & system
TDP575 W
Form factorPCIe dual-slot
PCIePCIe Gen5
MSRP$1,999

Compatible CUDA Toolkit versions

12.812.913.013.113.213.3
Rewritten 2026-09-01 using NVIDIA's official "NVIDIA RTX Blackwell GPU Architecture" whitepaper (images.nvidia.com/aem-dam/Solutions/geforce/ blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf) — its Appendix A ("Blackwell GB202 GPU") table gives a full RTX 3090/4090/5090 side-by-side comparison, read via pdftotext after WebFetch failed on this PDF's binary structure. Same source and methodology as the RTX 4090 entry; fp16TFLOPS/fp8TFLOPS (209.5/419 dense, FP32-accumulate) were already correct from an earlier pull and are confirmed unchanged here. tf32TFLOPS, bf16TFLOPS, fp16TFLOPS, and fp8TFLOPS are dense Tensor-Core figures using FP32 accumulate; the ...Sparse fields and int8TOPSSparse hold the structured-sparsity figures from the same rows. fp32TFLOPS (104.8) is the non-Tensor/vector "Peak FP32 TFLOPS" row — identical to the non-Tensor FP16/BF16 rows on Blackwell's unified shader cores, and to tf32TFLOPS dense for the same reason as the RTX 4090 entry. fp4TFLOPS (1,676 dense; 3,352 with sparsity, no dedicated sparse field in this schema) is new to Blackwell — this is also the exact source of the card's marketing "3,352 AI TOPS" figure (FP4-with-sparsity), the same way RTX 4090's "1,321 AI TOPS" turned out to be its FP8/INT8 sparse rate. int8TOPS/int8TOPSSparse (838/1,676) match the FP8-with-FP16- accumulate row exactly (same rate on this die). processNode ("TSMC 4N"), transistorCountBillion (92.2 — RTX 5090 uses the full, uncut GB202 die's transistor budget even though only 170 of 192 SMs are enabled), cacheMB (96, from "L2 Cache Size: 98,304 KB"), computeUnitCount (170 SMs), matrixCoreCount (680 Tensor Cores, 5th Gen), shaderCoreCount (21,760 CUDA cores), and peakClockMHz (2,407 MHz boost) are all read directly from the same whitepaper table. Die size (750 mm²) and RT Core count/TFLOPS are published but out of scope for this AI/ML-focused schema.

← All GPUs · Compare GPUs →