V100 SXM2 32GB

NVIDIAdatacenter
Overview
ArchitectureVolta (2017)
Launch dateMay 2017 (original 16GB V100); 32GB SXM2 variant added 2018
Compute capability7.0
Memory
VRAM32 GB HBM2
Memory bandwidth900 GB/s
ECCYes
Compute — vector
FP647.8 TFLOPS
FP3215.7 TFLOPS
Compute — matrix / tensor
FP16125 TFLOPS
Cores & clocks
Streaming Multiprocessors80
Shader cores5,120
Matrix / Tensor cores640
Board & system
TDP300 W
Form factorSXM2
InterconnectNVLink — 300 GB/s

Compatible CUDA Toolkit versions

10.010.110.211.211.311.611.711.812.012.112.212.312.412.512.612.812.9
Added 2026-09 to fill a gap this site's own README had flagged (oldest NVIDIA entry was 2020's A100; V100 is still a common ML training/inference GPU in the field). All figures sourced directly from NVIDIA's official V100 datasheet (images.nvidia.com/content/technologies/volta/pdf/ volta-v100-datasheet-update-us-1165301-r5.pdf, "V100 SXM2" column, footer dated Jan20), read via pdftotext after WebFetch could not parse this PDF's text layer directly (same issue noted on this site's A100/B300 entries). fp16TFLOPS (125) is the datasheet's single "Tensor Performance" row — Volta's 1st-gen Tensor Cores predate structured sparsity (introduced with Ampere), so there is no dense/sparse split and no fp16TFLOPSSparse value; this is the only Tensor-Core throughput row Volta publishes (no separate BF16/TF32/FP8/INT8 Tensor figures — BF16 and TF32 as Tensor-Core formats and FP8 didn't exist until later architectures). fp16TFLOPSVector (packed FP16 via CUDA cores) isn't in this datasheet either, so it's omitted rather than guessed at 2x fp32TFLOPS. This is the SXM2 form factor at 300W; the same datasheet lists a PCIe variant (V100 PCIe, 250W max power, 7/14/112 TFLOPS FP64/FP32/Tensor, 16 or 32GB, 900GB/s bandwidth, 32GB/sec PCIe Gen3 interconnect instead of NVLink) and a V100S PCIe refresh (8.2/16.4/130 TFLOPS, 32GB HBM2 only, 1134GB/s bandwidth, still 250W) — not modeled as separate entries here, noted for completeness. interconnectBandwidthGBs (300) is NVLink specifically; the SXM2 module also has a PCIe Gen3 host link not given a separate bandwidth figure in this table. No transistor count, process node, or MSRP in this datasheet — V100 shipped through OEM/server partners, not direct retail, and this table doesn't list die-level figures the way NVIDIA's longer Volta architecture whitepaper might. 2026-09-15 — added computeUnitCount (80 SMs) from NVIDIA's own Volta architecture whitepaper (WP-08608-001_v1.1), which states that "the Tesla V100 accelerator uses 80 SMs" of the full GV100's 84, and whose Table 1 lists the matching 5,120 FP32 cores and 640 Tensor Cores already recorded here (64 FP32 cores and 8 Tensor Cores per SM). Read via pdftotext after WebFetch failed on the PDF's binary structure — the same workaround this collection's RTX 4090 entry documents.

← All GPUs · Compare GPUs →