Instinct MI355X

AMDdatacenter
Overview
ArchitectureCDNA 4 (2025)
Launch date2025-06-12
Process nodeTSMC 3nm | 6nm FinFET
Transistor count185B
GPU target (gfx)gfx950
Memory
VRAM288 GB HBM3E
Memory bandwidth8,000 GB/s
Cache256 MB
ECCYes
Compute — vector
FP6478.6 TFLOPS
FP32157.3 TFLOPS
FP16157.3 TFLOPS
Compute — matrix / tensor
FP6478.6 TFLOPS
BF162,516.6 TFLOPS
FP16 (dense / sparse)2,516.6 / 5,033.2 TFLOPS
FP8 (dense / sparse)5,033.2 / 10,066.4 TFLOPS
FP6 (MXFP6)10,100 TFLOPS
FP4 (MXFP4 / NVFP4)10,100 TFLOPS
INT8 (dense / sparse)5,000 / 10,100 TOPS
Cores & clocks
Compute Units256
Shader cores16,384
Matrix / Tensor cores1,024
Peak clock2,400 MHz
Board & system
TDP1400 W
Form factorOAM Module
CoolingPassive & Active
PCIePCIe 5.0 x16
InterconnectInfinity Fabric — 153 GB/s
Multi-instance supportSR-IOV

Compatible ROCm versions

7.0.07.0.17.0.27.1.07.1.17.2.07.2.17.2.27.2.37.2.47.14.010.0.0
Same CDNA 4 die and gfx target as MI350X — MI355X is a liquid-cooled, higher-power/higher-clock variant (1,400W vs. 1,000W TBP), not a different chip. Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi350/mi355x.html, "Expand All" spec accordion), pulled 2026-09-01, cross-checked against the official MI355X brochure for the two headline dense/sparse figures. fp16TFLOPS/fp8TFLOPS (and their Sparse variants) use the brochure's precise decimal figures (2,516.6 / 5,033.2 / 5,033.2 / 10,066.4) rather than the product page's rounded PFLOP figures (2.5 / 5 / 5 / 10.1 PFLOPs) — same underlying numbers, brochure just carries one more significant digit. bf16TFLOPS is set equal to fp16TFLOPS: AMD's page lists identical PFLOP figures for "Peak BF16 Matrix" and "Peak FP16 Matrix" (both 2.5/5 PFLOPs dense/sparse), so the brochure's extra FP16 precision is assumed to carry over to BF16 as well (not separately confirmed at that precision). fp6TFLOPS/fp4TFLOPS (MXFP6/MXFP4, both 10.1 PFLOPs) and int8TOPS/int8TOPSSparse (5/10.1 POPs) come from the product page only, at its native 1-decimal-PFLOP precision — no brochure cross-check available for these. fp64TFLOPS and fp64TFLOPSMatrix are both 78.6 TFLOPs, and fp32TFLOPS (vector) equals the page's separately-listed "FP32 Matrix" figure (157.3) — AMD's page lists identical vector/matrix numbers for FP64 and FP32 on this part, unlike FP16 where vector (157.3) and matrix (2,500+) diverge sharply. interconnectBandwidthGBs (153) is the page's headline "Peak Infinity Fabric Link Bandwidth" figure across 7 links per GPU; the page also separately lists scale-up (153 GB/s) vs. scale-out (128 GB/s, 1 link) peak figures — 153 used here as the representative headline number. No msrpUSD — AMD doesn't publish Instinct list prices.

← All GPUs · Compare GPUs →