RTX 5090
NVIDIAconsumer| Overview | |
|---|---|
| Architecture | Blackwell (2025) |
| Process node | TSMC 4N |
| Transistor count | 92.2B |
| Compute capability | 12.0 |
| Memory | |
| VRAM | 32 GB GDDR7 |
| Memory bandwidth | 1,792 GB/s |
| Cache | 96 MB |
| Compute — vector | |
| FP32 | 104.8 TFLOPS |
| Compute — matrix / tensor | |
| TF32 | 104.8 TFLOPS |
| BF16 | 209.5 TFLOPS |
| FP16 (dense / sparse) | 209.5 / 419 TFLOPS |
| FP8 (dense / sparse) | 419 / 838 TFLOPS |
| FP4 (MXFP4 / NVFP4) | 1,676 TFLOPS |
| INT8 (dense / sparse) | 838 / 1,676 TOPS |
| Cores & clocks | |
| Streaming Multiprocessors | 170 |
| Shader cores | 21,760 |
| Matrix / Tensor cores | 680 |
| Peak clock | 2,407 MHz |
| Board & system | |
| TDP | 575 W |
| Form factor | PCIe dual-slot |
| PCIe | PCIe Gen5 |
| MSRP | $1,999 |
Compatible CUDA Toolkit versions
Rewritten 2026-09-01 using NVIDIA's official "NVIDIA RTX Blackwell GPU Architecture" whitepaper (images.nvidia.com/aem-dam/Solutions/geforce/ blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf) — its Appendix A ("Blackwell GB202 GPU") table gives a full RTX 3090/4090/5090 side-by-side comparison, read via pdftotext after WebFetch failed on this PDF's binary structure. Same source and methodology as the RTX 4090 entry; fp16TFLOPS/fp8TFLOPS (209.5/419 dense, FP32-accumulate) were already correct from an earlier pull and are confirmed unchanged here.
tf32TFLOPS, bf16TFLOPS, fp16TFLOPS, and fp8TFLOPS are dense Tensor-Core figures using FP32 accumulate; the ...Sparse fields and int8TOPSSparse hold the structured-sparsity figures from the same rows. fp32TFLOPS (104.8) is the non-Tensor/vector "Peak FP32 TFLOPS" row — identical to the non-Tensor FP16/BF16 rows on Blackwell's unified shader cores, and to tf32TFLOPS dense for the same reason as the RTX 4090 entry. fp4TFLOPS (1,676 dense; 3,352 with sparsity, no dedicated sparse field in this schema) is new to Blackwell — this is also the exact source of the card's marketing "3,352 AI TOPS" figure (FP4-with-sparsity), the same way RTX 4090's "1,321 AI TOPS" turned out to be its FP8/INT8 sparse rate. int8TOPS/int8TOPSSparse (838/1,676) match the FP8-with-FP16- accumulate row exactly (same rate on this die).
processNode ("TSMC 4N"), transistorCountBillion (92.2 — RTX 5090 uses the full, uncut GB202 die's transistor budget even though only 170 of 192 SMs are enabled), cacheMB (96, from "L2 Cache Size: 98,304 KB"), computeUnitCount (170 SMs), matrixCoreCount (680 Tensor Cores, 5th Gen), shaderCoreCount (21,760 CUDA cores), and peakClockMHz (2,407 MHz boost) are all read directly from the same whitepaper table. Die size (750 mm²) and RT Core count/TFLOPS are published but out of scope for this AI/ML-focused schema.