Instinct MI350X
AMDdatacenter| Overview | |
|---|---|
| Architecture | CDNA 4 (2025) |
| Launch date | 2025-06-12 |
| Process node | TSMC 3nm | 6nm FinFET |
| Transistor count | 185B |
| GPU target (gfx) | gfx950 |
| Memory | |
| VRAM | 288 GB HBM3E |
| Memory bandwidth | 8,000 GB/s |
| Cache | 256 MB |
| ECC | Yes |
| Compute — vector | |
| FP64 | 72.1 TFLOPS |
| FP32 | 144.2 TFLOPS |
| FP16 | 144.2 TFLOPS |
| Compute — matrix / tensor | |
| FP64 | 72.1 TFLOPS |
| BF16 | 2,309.6 TFLOPS |
| FP16 (dense / sparse) | 2,309.6 / 4,619.2 TFLOPS |
| FP8 (dense / sparse) | 4,614 / 9,227.4 TFLOPS |
| FP6 (MXFP6) | 9,200 TFLOPS |
| FP4 (MXFP4 / NVFP4) | 9,200 TFLOPS |
| INT8 (dense / sparse) | 4,600 / 9,200 TOPS |
| Cores & clocks | |
| Compute Units | 256 |
| Shader cores | 16,384 |
| Matrix / Tensor cores | 1,024 |
| Peak clock | 2,200 MHz |
| Board & system | |
| TDP | 1000 W |
| Form factor | OAM Module |
| Cooling | Passive OAM |
| PCIe | PCIe 5.0 x16 |
| Interconnect | Infinity Fabric — 153 GB/s |
| Multi-instance support | SR-IOV |
Compatible ROCm versions
Air-cooled sibling of MI355X — same CDNA4 die, gfx950 target, and 256 compute units/1,024 matrix cores, but a lower peak clock (2,200MHz vs. 2,400MHz) and TBP (1,000W vs. 1,400W liquid-cooled). Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi350/mi350x.html, "Expand All" spec accordion), pulled 2026-09-01.
fp16TFLOPS/fp8TFLOPS and their Sparse variants keep this file's previously-recorded brochure-precision decimal figures (2,309.6 / 4,619.2 / 4,614 / 9,227.4) rather than the product page's rounded PFLOP figures (2.3/4.6/4.6/9.2 PFLOPs) — same underlying numbers, brochure just carries one more significant digit. bf16TFLOPS is set equal to fp16TFLOPS on the same reasoning as MI355X: the page lists identical PFLOP figures for "Peak BF16 Matrix" and "Peak FP16 Matrix" (both 2.3/4.6 PFLOPs dense/sparse). fp6TFLOPS/fp4TFLOPS (MXFP6/MXFP4, both 9.2 PFLOPs) and int8TOPS/int8TOPSSparse (4.6/9.2 POPs) come from the product page only, at its native 1-decimal-PFLOP precision.
fp64TFLOPS/fp64TFLOPSMatrix are both 72.1 TFLOPs, and fp32TFLOPS (vector) equals the page's separately-listed "FP32 Matrix" figure (144.2) — identical vector/matrix numbers for FP64 and FP32 on this part, same pattern as MI355X.
interconnectBandwidthGBs (153) is the page's headline "Peak Infinity Fabric Link Bandwidth" across 7 links per GPU; scale-up/scale-out peak figures (153/128 GB/s) are also published but not separately modeled here. No msrpUSD — AMD doesn't publish Instinct list prices.