A100 SXM4 80GB
NVIDIAdatacenter| Overview | |
|---|---|
| Architecture | Ampere (2020) |
| Compute capability | 8.0 |
| Memory | |
| VRAM | 80 GB HBM2e |
| Memory bandwidth | 2,039 GB/s |
| Compute — vector | |
| FP64 | 9.7 TFLOPS |
| FP32 | 19.5 TFLOPS |
| Compute — matrix / tensor | |
| FP64 | 19.5 TFLOPS |
| TF32 | 156 TFLOPS |
| BF16 | 312 TFLOPS |
| FP16 (dense / sparse) | 312 / 624 TFLOPS |
| INT8 (dense / sparse) | 624 / 1,248 TOPS |
| Cores & clocks | |
| Streaming Multiprocessors | 108 |
| Shader cores | 6,912 |
| Matrix / Tensor cores | 432 |
| Board & system | |
| TDP | 400 W |
| Form factor | SXM |
| PCIe | PCIe Gen4 |
| Interconnect | NVLink — 600 GB/s |
| Multi-instance support | Up to 7 MIGs @ 10GB each |
Compatible CUDA Toolkit versions
fp8TFLOPS intentionally omitted — Ampere's Tensor Cores don't support native FP8; that was introduced with Hopper. All figures sourced directly from NVIDIA's official A100 datasheet (nvidia-a100-datasheet-nvidia-us-2188504-web.pdf, "80GB SXM" column), pulled 2026-09-01 via pdftotext extraction of the PDF (WebFetch could not parse this PDF's text layer directly).
fp64TFLOPS/fp64TFLOPSMatrix, fp32TFLOPS, tf32TFLOPS, bf16TFLOPS, fp16TFLOPS, and int8TOPS are all the vendor's own DENSE (non-sparse) figures — the datasheet lists each Tensor Core row as "X TFLOPS | Y TFLOPS*" with footnote "* With sparsity", so X is dense and Y is the sparsity figure. Only fp16 and int8 have a dedicated Sparse field in this schema (fp16TFLOPSSparse: 624, int8TOPSSparse: 1248); TF32's sparse figure is 312 TFLOPS and BF16's is 624 TFLOPS but there's no tf32/bf16 Sparse field — noted here instead.
tdpWatts (400) is the standard SXM4 config; the datasheet notes an HGX A100-80GB CTS (Custom Thermal Solution) SKU can go up to 500W. interconnectBandwidthGBs (600) is NVLink; the SXM module also exposes a PCIe Gen4 x16 host link at 64GB/s (pcieGen field records the generation, not this secondary bandwidth figure). multiInstanceSupport confirms MIG partitioning into up to 7 instances @ 10GB VRAM each. No msrpUSD — A100 ships through OEM/server partners, not direct retail. No transistor count, L2 cache size, or process node found in the datasheet itself (only in NVIDIA's longer Ampere architecture whitepaper, not cross-checked here) — omitted rather than sourced from a secondary site.
2026-09-15 — added computeUnitCount (108 SMs), shaderCoreCount (6,912 FP32 CUDA cores) and matrixCoreCount (432 third-generation Tensor Cores) from NVIDIA's own "NVIDIA Ampere Architecture In-Depth" developer blog post, which states the A100 product configuration of GA100 directly: "108 SMs; 64 FP32 CUDA Cores/SM, 6912 FP32 CUDA Cores per GPU; 4 third-generation Tensor Cores/SM, 432 third-generation Tensor Cores per GPU." These counts are identical across every A100 SKU (40GB/80GB, SXM4/PCIe) — only memory, bandwidth and power differ between them.