H200 SXM 141GB
NVIDIAdatacenter| Overview | |
|---|---|
| Architecture | Hopper (2024) |
| Compute capability | 9.0 |
| Memory | |
| VRAM | 141 GB HBM3e |
| Memory bandwidth | 4,800 GB/s |
| Compute — vector | |
| FP64 | 34 TFLOPS |
| FP32 | 67 TFLOPS |
| Compute — matrix / tensor | |
| FP64 | 67 TFLOPS |
| TF32 | 494.5 TFLOPS |
| BF16 | 989.5 TFLOPS |
| FP16 (dense / sparse) | 989.5 / 1,979 TFLOPS |
| FP8 (dense / sparse) | 1,979 / 3,958 TFLOPS |
| INT8 (dense / sparse) | 1,979 / 3,958 TOPS |
| Cores & clocks | |
| Streaming Multiprocessors | 132 |
| Shader cores | 16,896 |
| Matrix / Tensor cores | 528 |
| Board & system | |
| TDP | 700 W |
| Form factor | SXM |
| PCIe | PCIe Gen5 |
| Interconnect | NVLink — 900 GB/s |
| Multi-instance support | Up to 7 MIGs @ 18GB each |
Compatible CUDA Toolkit versions
Same Hopper compute die as H100 — H200 is a memory upgrade (more capacity, higher bandwidth), not a compute upgrade; compute figures are identical to the H100 SXM entry. Source: NVIDIA's official H200 product page (nvidia.com/en-us/data-center/h200/, "Specifications" table, H200 SXM column), pulled 2026-09-01. NVIDIA marks this table "Preliminary specifications. May be subject to change."
tf32TFLOPS, bf16TFLOPS, fp16TFLOPS, fp8TFLOPS, and int8TOPS are dense (non-sparse) figures, derived by halving NVIDIA's published "with sparsity" headline numbers (989/1,979/1,979/3,958/3,958), per the page's own footnote — same methodology as the H100 entry. The ...Sparse fields hold the vendor's directly published headline (sparse) figures. fp64 rows are not marked with the sparsity footnote, so fp64TFLOPS (34) and fp64TFLOPSMatrix (67) are used as-is.
multiInstanceSupport partition size (18GB) differs from H100's (10GB) to reflect H200's larger 141GB VRAM pool split across the same 7 MIG slices. No transistor count, L2 cache, or process node published on this page (SM / CUDA-core / Tensor-core counts have since been added from H200's shared GH100 die — see below). No msrpUSD — ships through OEM/server partners.
2026-09-15 — added computeUnitCount (132 SMs), shaderCoreCount (16,896 FP32 CUDA cores) and matrixCoreCount (528 fourth-generation Tensor Cores). NVIDIA does not publish these for H200 specifically: they are H100's figures, carried over because H200 reuses the same GH100 compute die unchanged (same 4N die, same 132 of 144 SMs enabled) and changes only the memory stack — HBM3e instead of HBM3, 141GB instead of 80GB. That is exactly why every compute figure in this file already matches this collection's H100 SXM5 entry line for line, with only the memory and MIG-partition rows differing. Recorded as a same-die inference, not a vendor-stated figure; if NVIDIA ever publishes a differing count for H200, these three fields are the ones to revisit.