Radeon Instinct MI60

AMDdatacenter
Overview
ArchitectureVega 7nm (GCN 5.1) (2018)
Launch dateNov 6, 2018 (announced; AMD said shipping to datacenter customers by end of 2018)
Process node7nm FinFET
Transistor count13.2B
GPU target (gfx)gfx906
Memory
VRAM32 GB HBM2
Memory bandwidth1,000 GB/s
ECCYes
Compute — vector
FP647.4 TFLOPS
FP3214.8 TFLOPS
FP1629.5 TFLOPS
Cores & clocks
Compute Units64
Shader cores4,096
Board & system
TDP300 W
PCIePCIe 4.0 x16
InterconnectInfinity Fabric Link (dual) — 200 GB/s
Multi-instance supportMxGPU SR-IOV (hardware virtualization, not MIG-style GPU partitioning)

ROCm: runs, but not officially supported

AMD has never named the Radeon Instinct MI60 as a supported GPU. These ROCm releases support its gfx906 target through other cards, so it generally runs without any workaround, but AMD doesn't validate it:

4.5.04.5.25.0.05.0.15.0.25.1.05.1.15.1.35.2.05.2.15.2.35.3.05.3.25.3.35.4.05.4.15.4.25.4.35.5.05.5.15.6.05.6.15.7.05.7.16.0.06.0.26.1.06.1.16.1.26.1.56.2.06.2.16.2.26.2.46.3.06.3.16.3.26.3.3

Find the PyTorch build to install on this card →

Libraries on the Radeon Instinct MI60

Whether each library's own requirements cover this card. Each row links to the version and source it was checked against on library support.

LibrarySupportedDetails
vLLM 0.30.0No
FlashAttention-2 2.8.3.post1No
FlashAttention-3 main (2026-09-27)No
bitsandbytes 0.50.2No
TensorRT 11.3.0No
llama.cpp master (2026-09-27)Partlyneeds ROCm 6.3.3 or older for the ROCm backend; otherwise Vulkan backend only

Native low-precision formats

Whether the card's matrix hardware runs these formats natively. Without native support a format can still work, but through slower emulation or conversion, not at the card's rated speed.

BF16FP8FP4
Not documentedNot documentedNot documented

Source: AMD ROCm 10.0.0 data types and precision support (matrix core tables); this architecture isn't in AMD's table.

Added 2026-09 — the full-die sibling of the MI50 (same Vega 20 / gfx906 silicon), and a card that comes up constantly in local-LLM discussions alongside it. Shares the MI50's ROCm history exactly, since ROCm grants support per gfx target: see the MI50 and Radeon VII entries. Figures sourced from AMD's official launch press release of 2018-11-06 ("AMD Unveils World's First 7nm Datacenter GPUs", as distributed via GlobeNewswire): 29.5 TFLOPS FP16, 14.8 TFLOPS FP32, 7.4 TFLOPS FP64 peak theoretical, 13.2 billion transistors on a 331.46mm² die at 300W (footnote 1); 32GB HBM2 with full-chip ECC covering "HBM2 memory and internal GPU structures" (footnote 6); dual Infinity Fabric Links at "up to 200 GB/s peak theoretical GPU to GPU" per card, 264 GB/s aggregate with PCIe Gen 4 (footnote 4); MxGPU hardware virtualization. memoryBandwidthGBs (1000) follows the release's "up to 1 TB/s" wording, matching how the MI50 entry records the same memory subsystem. AMD's MI60 datasheet PDF (amd.com/system/files/documents/ radeon-instinct-mi60-datasheet.pdf) could not be retrieved at time of writing. computeUnitCount (64) and shaderCoreCount (4096) are not stated in the press release: they are the full Vega 20 die (64 stream processors per CU; the MI50's datasheet lists 60 CUs / 3840 for the cut-down part), and are consistent with AMD's own FP32 figure — 4096 x 2 FLOP x 1.8 GHz = 14.7 TFLOPS. No INT8 figure is given in the press release, so int8TOPS is omitted rather than inferred. fp16TFLOPSVector is Vega's Rapid-Packed-Math rate on the stream processors: gfx906 predates AMD's Matrix Cores (introduced with CDNA/MI100).

← All GPUs · Compare GPUs →