Instinct MI455X
AMDdatacenter| Overview | |
|---|---|
| Architecture | CDNA 5 (2026) |
| Launch date | 2026-07-23 |
| Process node | TSMC 2nm | 3nm FinFET |
| Transistor count | 320B |
| Memory | |
| VRAM | 432 GB HBM4 |
| Memory bandwidth | 23,300 GB/s |
| Cache | 192 MB |
| ECC | Yes |
| Compute — vector | |
| FP64 | 5 TFLOPS |
| FP32 | 315 TFLOPS |
| FP16 | 315 TFLOPS |
| Compute — matrix / tensor | |
| FP64 | 5 TFLOPS |
| BF16 | 5,000 TFLOPS |
| FP16 (dense / sparse) | 5,000 / 10,100 TFLOPS |
| FP8 (dense / sparse) | 10,050 / 20,100 TFLOPS |
| FP6 (MXFP6) | 20,100 TFLOPS |
| FP4 (MXFP4 / NVFP4) | 40,300 TFLOPS |
| INT8 (dense / sparse) | 5,000 / 10,100 TOPS |
| Cores & clocks | |
| Work Group Processors | 256 |
| Peak clock | 2,400 MHz |
| Board & system | |
| Form factor | Enhanced Accelerator Module (EAM) |
| Cooling | Direct Liquid Cooling (DLC) |
| Interconnect | UALink / UALoE — 3,600 GB/s |
| Multi-instance support | SR-IOV |
Part of the MI400 series / "Helios" rack-scale platform. AMD's product page (amd.com/en/products/accelerators/instinct/mi400/mi455x.html, "Expand All" spec accordion) now publishes a full spec sheet as of this pull (2026-09-01) — this collection's earlier entry, written before launch, had to derive/estimate several figures that are now directly confirmed.
fp16TFLOPS/fp16TFLOPSSparse (5,000/10,100, "Peak Matrix FP16 Performance") and bf16TFLOPS (5,000, "Peak bfloat16 (BF16) Matrix Performance" — its 10,100 sparse variant isn't captured in a separate field) are both directly stated and identical, unlike CDNA4 where BF16 had to be inferred equal to FP16. fp8TFLOPS/fp8TFLOPSSparse are still DERIVED, not directly stated: the page lists a single "Peak OCP FP8 Performance: 20.1 PFLOPs" with no dense/sparse split — by analogy to the FP16/INT8 rows on this same page (both showing dense = 2x lower than their own sparse figure, and dense FP8 = 2x dense FP16 on every CDNA3/CDNA4 part in this collection), 20.1 PFLOPs is treated as the SPARSE figure, giving a derived dense fp8TFLOPS of ~10,050 — same estimate this file used pre-launch, now more confidently cross-checked against the newly-published FP16/INT8 dense:sparse ratios. fp6TFLOPS (MXFP6, 20,100) and fp4TFLOPS (MXFP4, 40,300) are used as directly published, single-figure, precision unconfirmed as dense or sparse.
fp64 is surprisingly low on this part — both vector and matrix rated at just 5 TFLOPs (vs. 78.6 on MI355X) — taken directly from the page as published, not adjusted; MI455X is evidently not built for FP64/HPC workloads the way CDNA3/4 parts were. fp32 vector and matrix are both 315 TFLOPs (also equal to the page's separately-stated FP16 vector figure).
computeUnitLabel is "Work Group Processors" (256) — AMD's new CDNA5 terminology replacing "Compute Units"; no separate Stream Processor/Matrix Core counts are published for this part. interconnectBandwidthGBs (3,600) uses the page's "Scale-up (Peak) UALoE Bi-directional Bandwidth" figure as the headline number; the complementary "Scale-out (Peak) UALink Bi-directional Bandwidth" (600 GB/s) is a separate, smaller network-fabric figure not captured in its own field.
tdpWatts and targetId are still both omitted: no TBP or PCIe bus type is published on this page (unofficial estimates put TBP "north of 2kW" based on Helios rack power budgets — not used here), and ROCm/CDNA5 support isn't reflected in this site's `rocm` collection yet, so there's no confirmed gfx target to show. No msrpUSD — sold as part of rack-scale deals, not a per-unit list price.