Instinct MI355X
AMDdatacenter| Overview | |
|---|---|
| Architecture | CDNA 4 (2025) |
| Launch date | 2025-06-12 |
| Process node | TSMC 3nm | 6nm FinFET |
| Transistor count | 185B |
| GPU target (gfx) | gfx950 |
| Memory | |
| VRAM | 288 GB HBM3E |
| Memory bandwidth | 8,000 GB/s |
| Cache | 256 MB |
| ECC | Yes |
| Compute — vector | |
| FP64 | 78.6 TFLOPS |
| FP32 | 157.3 TFLOPS |
| FP16 | 157.3 TFLOPS |
| Compute — matrix / tensor | |
| FP64 | 78.6 TFLOPS |
| BF16 | 2,516.6 TFLOPS |
| FP16 (dense / sparse) | 2,516.6 / 5,033.2 TFLOPS |
| FP8 (dense / sparse) | 5,033.2 / 10,066.4 TFLOPS |
| FP6 (MXFP6) | 10,100 TFLOPS |
| FP4 (MXFP4 / NVFP4) | 10,100 TFLOPS |
| INT8 (dense / sparse) | 5,000 / 10,100 TOPS |
| Cores & clocks | |
| Compute Units | 256 |
| Shader cores | 16,384 |
| Matrix / Tensor cores | 1,024 |
| Peak clock | 2,400 MHz |
| Board & system | |
| TDP | 1400 W |
| Form factor | OAM Module |
| Cooling | Passive & Active |
| PCIe | PCIe 5.0 x16 |
| Interconnect | Infinity Fabric — 153 GB/s |
| Multi-instance support | SR-IOV |
Compatible ROCm versions
Same CDNA 4 die and gfx target as MI350X — MI355X is a liquid-cooled, higher-power/higher-clock variant (1,400W vs. 1,000W TBP), not a different chip. Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi350/mi355x.html, "Expand All" spec accordion), pulled 2026-09-01, cross-checked against the official MI355X brochure for the two headline dense/sparse figures.
fp16TFLOPS/fp8TFLOPS (and their Sparse variants) use the brochure's precise decimal figures (2,516.6 / 5,033.2 / 5,033.2 / 10,066.4) rather than the product page's rounded PFLOP figures (2.5 / 5 / 5 / 10.1 PFLOPs) — same underlying numbers, brochure just carries one more significant digit. bf16TFLOPS is set equal to fp16TFLOPS: AMD's page lists identical PFLOP figures for "Peak BF16 Matrix" and "Peak FP16 Matrix" (both 2.5/5 PFLOPs dense/sparse), so the brochure's extra FP16 precision is assumed to carry over to BF16 as well (not separately confirmed at that precision). fp6TFLOPS/fp4TFLOPS (MXFP6/MXFP4, both 10.1 PFLOPs) and int8TOPS/int8TOPSSparse (5/10.1 POPs) come from the product page only, at its native 1-decimal-PFLOP precision — no brochure cross-check available for these.
fp64TFLOPS and fp64TFLOPSMatrix are both 78.6 TFLOPs, and fp32TFLOPS (vector) equals the page's separately-listed "FP32 Matrix" figure (157.3) — AMD's page lists identical vector/matrix numbers for FP64 and FP32 on this part, unlike FP16 where vector (157.3) and matrix (2,500+) diverge sharply.
interconnectBandwidthGBs (153) is the page's headline "Peak Infinity Fabric Link Bandwidth" figure across 7 links per GPU; the page also separately lists scale-up (153 GB/s) vs. scale-out (128 GB/s, 1 link) peak figures — 153 used here as the representative headline number.
No msrpUSD — AMD doesn't publish Instinct list prices.