{
  "dataset": "gpus",
  "source": "https://gpucompat.com/data/gpus.json",
  "about": "https://gpucompat.com/about-the-data/",
  "count": 65,
  "entries": [
    {
      "id": "amd-instinct-mi100",
      "vendor": "amd",
      "name": "Instinct MI100",
      "category": "datacenter",
      "architecture": "CDNA 1",
      "releaseYear": 2020,
      "launchDate": "2020-11-16",
      "targetId": "gfx908",
      "vramGB": 32,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 1200,
      "eccSupport": true,
      "fp64TFLOPS": 11.5,
      "fp32TFLOPS": 23.1,
      "bf16TFLOPS": 92.3,
      "fp16TFLOPS": 184.6,
      "int8TOPS": 92.3,
      "int4TOPS": 92.3,
      "processNode": "TSMC 7nm FinFET",
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 120,
      "shaderCoreCount": 7680,
      "peakClockMHz": 1502,
      "tdpWatts": 300,
      "formFactor": "PCIe Add-in Card",
      "coolingType": "Passive",
      "pcieGen": "PCIe 4.0 x16, PCIe 3.0 x16",
      "interconnect": "Infinity Fabric",
      "interconnectBandwidthGBs": 92,
      "multiInstanceSupport": "N/A",
      "notes": "fp8TFLOPS omitted — CDNA 1 predates FP8 matrix support (added in CDNA 3 with MI300). No transistor count or Matrix Cores count is published for this generation. Unlike MI250X/MI300-series, bfloat16 (92.3 TFLOPs) is HALF of fp16TFLOPS (184.6) on this part rather than equal to it — AMD's page lists them as genuinely different rates here; int8TOPS/int4TOPS (92.3 each) match the bf16 rate, not the fp16 rate. No separate FP64/FP32 matrix figures are published — only the single \"Peak Double/Single Precision (FP64/FP32) Performance\" rows, used directly (FP32 Matrix and FP32 vector figures happen to be identical on this page, 23.1, so only one field is populated). No msrpUSD — AMD doesn't publish Instinct list prices. multiInstanceSupport set to N/A — no SR-IOV/partitioning listed (introduced starting with MI300-series). Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi100.html, \"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-instinct-mi210",
      "vendor": "amd",
      "name": "Instinct MI210",
      "category": "datacenter",
      "architecture": "CDNA 2",
      "releaseYear": 2022,
      "targetId": "gfx90a",
      "vramGB": 64,
      "memoryType": "HBM2e",
      "memoryBandwidthGBs": 1600,
      "eccSupport": true,
      "fp64TFLOPS": 22.6,
      "fp32TFLOPS": 22.6,
      "fp64TFLOPSMatrix": 45.3,
      "bf16TFLOPS": 181,
      "fp16TFLOPS": 181,
      "int8TOPS": 181,
      "processNode": "TSMC 6nm FinFET",
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 104,
      "shaderCoreCount": 6656,
      "peakClockMHz": 1700,
      "tdpWatts": 300,
      "formFactor": "PCIe dual-slot FHFL",
      "pcieGen": "PCIe 4.0 x16",
      "interconnect": "Infinity Fabric (3rd gen)",
      "multiInstanceSupport": "N/A",
      "notes": "Added 2026-09-06 during a `gpus` coverage pass -- a single-GCD, PCIe add-in-card sibling of the already-tracked MI250X (same CDNA 2 generation, same gfx90a target), commonly deployed as a single-GPU on-prem/workstation ROCm accelerator where MI250X's dual-die OAM module isn't practical.\nAMD's own product PDFs (instinct-mi210-brochure.pdf, amd-instinct-gpu-family-brochure.pdf) repeatedly timed out on direct fetch -- unlike most entries in this collection, these figures come from a secondary spec aggregator (flopper.io) and a semiconductor-analysis writeup (chipsandcheese.com's \"AMD's Radeon Instinct MI210: GCN Lives On\"), cross-checked for internal consistency against this site's already-verified MI250X entry rather than against AMD directly: fp64TFLOPSMatrix (45.3) is almost exactly 2x fp64TFLOPS (22.6), the same vector:matrix ratio MI250X's page publishes explicitly (47.9 -> 95.7); vramGB (64) and memoryBandwidthGBs (1600) are almost exactly half of MI250X's dual-GCD board-level figures (128 / 3276.8), consistent with MI210 being a single-GCD product built from the same die.\ncomputeUnitCount (104) and shaderCoreCount (6656, derived from 104 CUs x 64 stream processors/CU -- the same ratio MI250X's page confirms directly, 14080/220=64) come from chipsandcheese specifically, not flopper.io; 104 CUs is noticeably less than a literal half of MI250X's 220 (110), which is plausible binning/yield variance between the two products rather than an error, but is flagged here since it's the one figure that doesn't cleanly reconcile.\nint4TOPS and interconnectBandwidthGBs are omitted -- no MI210-specific figure was found in any source consulted, and int4TOPS on MI250X (383, identical to its int8TOPS) isn't confirmed to hold for MI210 specifically rather than assumed. No msrpUSD -- AMD doesn't publish Instinct list prices, matching every other Instinct entry in this collection.\n"
    },
    {
      "id": "amd-instinct-mi250x",
      "vendor": "amd",
      "name": "Instinct MI250X",
      "category": "datacenter",
      "architecture": "CDNA 2",
      "releaseYear": 2021,
      "launchDate": "2021-11-08",
      "targetId": "gfx90a",
      "vramGB": 128,
      "memoryType": "HBM2e",
      "memoryBandwidthGBs": 3276.8,
      "eccSupport": true,
      "fp64TFLOPS": 47.9,
      "fp32TFLOPS": 47.9,
      "fp64TFLOPSMatrix": 95.7,
      "bf16TFLOPS": 383,
      "fp16TFLOPS": 383,
      "int8TOPS": 383,
      "int4TOPS": 383,
      "processNode": "TSMC 6nm FinFET",
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 220,
      "shaderCoreCount": 14080,
      "peakClockMHz": 1700,
      "tdpWatts": 500,
      "formFactor": "OAM Module",
      "coolingType": "Passive OAM",
      "pcieGen": "PCIe 4.0 x16",
      "interconnect": "Infinity Fabric (3rd gen)",
      "interconnectBandwidthGBs": 100,
      "multiInstanceSupport": "N/A",
      "notes": "fp8TFLOPS omitted — CDNA 2 has no native FP8 matrix support (added in CDNA 3 with MI300). No \"Matrix Cores\" count or transistor count is published for this generation (AMD started listing those with MI300/CDNA3) — matrixCoreCount intentionally omitted rather than guessed. MI250X is a dual-die (OAM) module; the page lists TDP as \"500W | 560W Peak\" — 500W (typical) kept as tdpWatts, matching this file's prior figure. fp32TFLOPS uses the plain (non-matrix) \"Peak Single Precision (FP32) Performance\" figure (47.9) for consistency with how this collection reports vector FP32 elsewhere; the page's separate \"FP32 Matrix\" figure (95.7, 2x vector) is not captured in its own field for this older generation. fp64TFLOPS (47.9, vector) and fp64TFLOPSMatrix (95.7, 2x) do differ, matching the FP32 pattern. bfloat16 and FP16 are published as identical figures (383 TFLOPs), used for both bf16TFLOPS and fp16TFLOPS. multiInstanceSupport set to N/A — no SR-IOV/MIG-equivalent partitioning is listed on this page (introduced starting with MI300-series). Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi200/mi250x.html, \"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-instinct-mi300a",
      "vendor": "amd",
      "name": "Instinct MI300A",
      "category": "datacenter",
      "architecture": "CDNA 3",
      "releaseYear": 2023,
      "targetId": "gfx942",
      "vramGB": 128,
      "memoryType": "HBM3 (unified CPU+GPU)",
      "memoryBandwidthGBs": 5300,
      "cacheMB": 256,
      "eccSupport": true,
      "fp64TFLOPS": 61.3,
      "fp32TFLOPS": 122.6,
      "fp64TFLOPSMatrix": 122.6,
      "tf32TFLOPS": 490.3,
      "bf16TFLOPS": 980.6,
      "fp16TFLOPS": 980.6,
      "fp8TFLOPS": 1961.2,
      "int8TOPS": 1961,
      "fp16TFLOPSSparse": 1961.2,
      "fp8TFLOPSSparse": 3922.4,
      "processNode": "TSMC 5nm | 6nm FinFET",
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 228,
      "shaderCoreCount": 14592,
      "matrixCoreCount": 912,
      "peakClockMHz": 2100,
      "tdpWatts": 550,
      "formFactor": "SH5 socket (APU, not an OAM/PCIe card)",
      "notes": "Added 2026-09-06 during a `gpus` coverage pass -- an APU sibling of the already-tracked MI300X (same CDNA 3 die family, same gfx942 target) that integrates 24 Zen 4 x86 CPU cores alongside the GPU chiplets on one package, sharing one pool of HBM3 as unified CPU+GPU memory (hence 128GB here vs. MI300X's 192GB -- some HBM/die budget goes to the CPU side instead). This CPU core count isn't captured in any field of this GPU-focused schema; noted here since it's the entire point of this SKU and not otherwise visible on this page.\nAMD's own MI300A datasheet PDF timed out on direct fetch (same issue as MI210 above); figures instead come from secondary aggregators (gpucost.org, cputronic.com), cross-checked by linear-scaling from this site's already- verified MI300X entry using the 228/304 = 0.75 ratio of MI300A's to MI300X's compute-unit count (both run at the same 2100MHz peak clock): 0.75 x MI300X's fp32TFLOPS (163.4) = 122.55 -> matches the sourced 122.6 almost exactly; 0.75 x MI300X's fp64TFLOPS/fp64TFLOPSMatrix (81.7/163.4) = 61.275/122.55 -> matches 61.3/122.6; 0.75 x MI300X's tf32TFLOPS (653.7) = 490.275 -> matches 490.3. matrixCoreCount (912) and shaderCoreCount (14,592) both reduce to the exact same per-CU ratios MI300X's page confirms directly (4 matrix cores and 64 stream processors per CU).\nThe secondary sources label several of these as \"with sparsity\" figures, but the numbers themselves match MI300X's DENSE figures scaled by 0.75, not its sparse figures (which would be roughly double) -- treated as a mislabeling in those sources rather than a real dense/sparse difference, so fp16TFLOPS/fp8TFLOPS above are recorded as dense, with fp16TFLOPSSparse/fp8TFLOPSSparse then derived by doubling, matching the dense:sparse ratio already established on the MI300X entry. int8TOPS (1,961) reuses the fp8TFLOPS dense figure -- no MI300A-specific INT8 number distinct from FP8 was found in any source consulted (MI300X's own page does publish a slightly different INT8 figure from its FP8 one, 2,600 vs. 2,614.9, so this may be an oversimplification, not a confirmed match).\ntdpWatts (550) is AMD's stated air/liquid-cooled maximum; a 760W figure also appears in the same sources for a liquid-cooled-only configuration, not used here to match this collection's convention elsewhere (MI250X) of using the lower/typical figure when two are published. No pcieGen, interconnectBandwidthGBs, multiInstanceSupport, transistorCountBillion, or msrpUSD -- none were found specifically confirmed for this SKU (as opposed to assumed identical to MI300X), so left unset rather than guessed.\n"
    },
    {
      "id": "amd-instinct-mi300x",
      "vendor": "amd",
      "name": "Instinct MI300X",
      "category": "datacenter",
      "architecture": "CDNA 3",
      "releaseYear": 2023,
      "launchDate": "2023-12-06",
      "targetId": "gfx942",
      "vramGB": 192,
      "memoryType": "HBM3",
      "memoryBandwidthGBs": 5325,
      "cacheMB": 256,
      "eccSupport": true,
      "fp64TFLOPS": 81.7,
      "fp32TFLOPS": 163.4,
      "fp64TFLOPSMatrix": 163.4,
      "tf32TFLOPS": 653.7,
      "bf16TFLOPS": 1307.4,
      "fp16TFLOPS": 1307.4,
      "fp8TFLOPS": 2614.9,
      "int8TOPS": 2600,
      "fp16TFLOPSSparse": 2614.9,
      "fp8TFLOPSSparse": 5229.8,
      "int8TOPSSparse": 5220,
      "processNode": "TSMC 5nm | 6nm FinFET",
      "transistorCountBillion": 153,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 304,
      "shaderCoreCount": 19456,
      "matrixCoreCount": 1216,
      "peakClockMHz": 2100,
      "tdpWatts": 750,
      "formFactor": "OAM Module",
      "coolingType": "Passive OAM",
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "Infinity Fabric",
      "interconnectBandwidthGBs": 128,
      "multiInstanceSupport": "SR-IOV",
      "notes": "Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi300/mi300x.html, \"Expand All\" spec accordion), pulled 2026-09-01, cross-checked against this file's previously-recorded MI300X/MI325X datasheet decimal figures.\nfp16TFLOPS/fp8TFLOPS and their Sparse variants keep the earlier brochure-precision figures (1,307.4 / 2,614.9 / 2,614.9 / 5,229.8) rather than the product page's rounded PFLOP figures (1.3/2.61/2.61/5.22 PFLOPs) — same underlying numbers, brochure carries one more significant digit. Unlike the CDNA4 (MI350-series) pages, MI300X's page publishes only ONE FP16 figure (\"Peak Half Precision (FP16) Performance\", no separate \"Matrix\" vs. vector split) — treated as the matrix/tensor rate here (fp16TFLOPSVector intentionally omitted, not zero) since 1.3 PFLOPs is far above FP32 vector throughput and can only be a tensor-core rate.\nfp64TFLOPS (81.7, vector) and fp64TFLOPSMatrix (163.4) genuinely DIFFER on this part — unlike MI350X/MI355X where AMD lists identical vector/matrix FP64 figures, MI300X's matrix engines run FP64 at exactly 2x the vector rate. fp32TFLOPS uses the page's \"FP32 Matrix\" and plain \"FP32\" figures, which are identical (163.4) on this part.\ntf32TFLOPS (653.7, dense) is new to this collection — MI300X/CDNA3 publishes a TF32 matrix rate (AMD's answer to NVIDIA's Tensor Float 32) that CDNA4 (MI350-series) no longer lists; its structured-sparsity variant (~1.3 PFLOPs, ~2x) isn't captured in a separate field. bf16TFLOPS is set equal to fp16TFLOPS: the page's \"Peak bfloat16\" figure (1.3 PFLOPs) matches FP16 exactly, same pattern as every other Instinct part in this collection. int8TOPS/int8TOPSSparse (2,600/5,220 POPs) are the page's own values at their native precision — no separate brochure decimal available.\ntdpWatts (750) is labelled \"Typical Board Power (TBP)\" on the page but the value itself reads \"750W Peak\" — kept as-is, same figure this collection has always used for MI300X. interconnectBandwidthGBs (128) is the page's single \"Peak Infinity Fabric Link Bandwidth\" figure across 8 links; MI300X's page (unlike MI350-series) doesn't publish a separate scale-up/scale-out breakdown. No msrpUSD — AMD doesn't publish Instinct list prices; unofficial street-price estimates (~$15,000) are deliberately not used.\n"
    },
    {
      "id": "amd-instinct-mi325x",
      "vendor": "amd",
      "name": "Instinct MI325X",
      "category": "datacenter",
      "architecture": "CDNA 3",
      "releaseYear": 2024,
      "launchDate": "2024-10-10",
      "targetId": "gfx942",
      "vramGB": 256,
      "memoryType": "HBM3E",
      "memoryBandwidthGBs": 6000,
      "cacheMB": 256,
      "eccSupport": true,
      "fp64TFLOPS": 81.7,
      "fp32TFLOPS": 163.4,
      "fp64TFLOPSMatrix": 163.4,
      "tf32TFLOPS": 653.7,
      "bf16TFLOPS": 1307.4,
      "fp16TFLOPS": 1307.4,
      "fp8TFLOPS": 2614.9,
      "int8TOPS": 2600,
      "fp16TFLOPSSparse": 2614.9,
      "fp8TFLOPSSparse": 5229.8,
      "int8TOPSSparse": 5220,
      "processNode": "TSMC 5nm | 6nm FinFET",
      "transistorCountBillion": 153,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 304,
      "shaderCoreCount": 19456,
      "matrixCoreCount": 1216,
      "peakClockMHz": 2100,
      "tdpWatts": 1000,
      "formFactor": "OAM Module",
      "coolingType": "Passive OAM",
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "Infinity Fabric",
      "interconnectBandwidthGBs": 128,
      "multiInstanceSupport": "SR-IOV",
      "notes": "Same CDNA3 compute die as MI300X (identical compute-unit/matrix-core counts and dense TFLOPS figures) — MI325X is a memory upgrade (192GB→256GB, HBM3→HBM3E, 5.3TB/s→6TB/s bandwidth) at a higher power envelope (1,000W vs. 750W). Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi300/mi325x.html, \"Expand All\" spec accordion), pulled 2026-09-01.\nAll compute-throughput fields (fp64/fp32/tf32/bf16/fp16/fp8/int8 and their Matrix/Sparse variants) mirror MI300X's file exactly — same sourcing caveats apply: fp16TFLOPS/fp8TFLOPS keep this collection's brochure-precision decimals (1,307.4/2,614.9, sparse 2,614.9/5,229.8) over the page's rounded PFLOP figures; fp64 vector (81.7) and matrix (163.4) genuinely differ 2x on this die; tf32TFLOPS (653.7, dense) has an unmodeled ~2x sparse variant; bf16TFLOPS is set equal to fp16TFLOPS per the page's identical PFLOP figures; int8TOPS/int8TOPSSparse (2,600/5,220 POPs) are page-native precision only.\ntdpWatts (1000) is labelled \"Typical Board Power (TBP)\" on the page but the value itself reads \"1000W Peak\". No msrpUSD — AMD doesn't publish Instinct list prices.\n"
    },
    {
      "id": "amd-instinct-mi350x",
      "vendor": "amd",
      "name": "Instinct MI350X",
      "category": "datacenter",
      "architecture": "CDNA 4",
      "releaseYear": 2025,
      "launchDate": "2025-06-12",
      "targetId": "gfx950",
      "vramGB": 288,
      "memoryType": "HBM3E",
      "memoryBandwidthGBs": 8000,
      "cacheMB": 256,
      "eccSupport": true,
      "fp64TFLOPS": 72.1,
      "fp32TFLOPS": 144.2,
      "fp16TFLOPSVector": 144.2,
      "fp64TFLOPSMatrix": 72.1,
      "bf16TFLOPS": 2309.6,
      "fp16TFLOPS": 2309.6,
      "fp8TFLOPS": 4614,
      "fp6TFLOPS": 9200,
      "fp4TFLOPS": 9200,
      "int8TOPS": 4600,
      "fp16TFLOPSSparse": 4619.2,
      "fp8TFLOPSSparse": 9227.4,
      "int8TOPSSparse": 9200,
      "processNode": "TSMC 3nm | 6nm FinFET",
      "transistorCountBillion": 185,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 256,
      "shaderCoreCount": 16384,
      "matrixCoreCount": 1024,
      "peakClockMHz": 2200,
      "tdpWatts": 1000,
      "formFactor": "OAM Module",
      "coolingType": "Passive OAM",
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "Infinity Fabric",
      "interconnectBandwidthGBs": 153,
      "multiInstanceSupport": "SR-IOV",
      "notes": "Air-cooled sibling of MI355X — same CDNA4 die, gfx950 target, and 256 compute units/1,024 matrix cores, but a lower peak clock (2,200MHz vs. 2,400MHz) and TBP (1,000W vs. 1,400W liquid-cooled). Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi350/mi350x.html, \"Expand All\" spec accordion), pulled 2026-09-01.\nfp16TFLOPS/fp8TFLOPS and their Sparse variants keep this file's previously-recorded brochure-precision decimal figures (2,309.6 / 4,619.2 / 4,614 / 9,227.4) rather than the product page's rounded PFLOP figures (2.3/4.6/4.6/9.2 PFLOPs) — same underlying numbers, brochure just carries one more significant digit. bf16TFLOPS is set equal to fp16TFLOPS on the same reasoning as MI355X: the page lists identical PFLOP figures for \"Peak BF16 Matrix\" and \"Peak FP16 Matrix\" (both 2.3/4.6 PFLOPs dense/sparse). fp6TFLOPS/fp4TFLOPS (MXFP6/MXFP4, both 9.2 PFLOPs) and int8TOPS/int8TOPSSparse (4.6/9.2 POPs) come from the product page only, at its native 1-decimal-PFLOP precision.\nfp64TFLOPS/fp64TFLOPSMatrix are both 72.1 TFLOPs, and fp32TFLOPS (vector) equals the page's separately-listed \"FP32 Matrix\" figure (144.2) — identical vector/matrix numbers for FP64 and FP32 on this part, same pattern as MI355X.\ninterconnectBandwidthGBs (153) is the page's headline \"Peak Infinity Fabric Link Bandwidth\" across 7 links per GPU; scale-up/scale-out peak figures (153/128 GB/s) are also published but not separately modeled here. No msrpUSD — AMD doesn't publish Instinct list prices.\n"
    },
    {
      "id": "amd-instinct-mi355x",
      "vendor": "amd",
      "name": "Instinct MI355X",
      "category": "datacenter",
      "architecture": "CDNA 4",
      "releaseYear": 2025,
      "launchDate": "2025-06-12",
      "targetId": "gfx950",
      "vramGB": 288,
      "memoryType": "HBM3E",
      "memoryBandwidthGBs": 8000,
      "cacheMB": 256,
      "eccSupport": true,
      "fp64TFLOPS": 78.6,
      "fp32TFLOPS": 157.3,
      "fp16TFLOPSVector": 157.3,
      "fp64TFLOPSMatrix": 78.6,
      "bf16TFLOPS": 2516.6,
      "fp16TFLOPS": 2516.6,
      "fp8TFLOPS": 5033.2,
      "fp6TFLOPS": 10100,
      "fp4TFLOPS": 10100,
      "int8TOPS": 5000,
      "fp16TFLOPSSparse": 5033.2,
      "fp8TFLOPSSparse": 10066.4,
      "int8TOPSSparse": 10100,
      "processNode": "TSMC 3nm | 6nm FinFET",
      "transistorCountBillion": 185,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 256,
      "shaderCoreCount": 16384,
      "matrixCoreCount": 1024,
      "peakClockMHz": 2400,
      "tdpWatts": 1400,
      "formFactor": "OAM Module",
      "coolingType": "Passive & Active",
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "Infinity Fabric",
      "interconnectBandwidthGBs": 153,
      "multiInstanceSupport": "SR-IOV",
      "notes": "Same CDNA 4 die and gfx target as MI350X — MI355X is a liquid-cooled, higher-power/higher-clock variant (1,400W vs. 1,000W TBP), not a different chip. Source: AMD's official product/spec page (amd.com/en/products/accelerators/instinct/mi350/mi355x.html, \"Expand All\" spec accordion), pulled 2026-09-01, cross-checked against the official MI355X brochure for the two headline dense/sparse figures.\nfp16TFLOPS/fp8TFLOPS (and their Sparse variants) use the brochure's precise decimal figures (2,516.6 / 5,033.2 / 5,033.2 / 10,066.4) rather than the product page's rounded PFLOP figures (2.5 / 5 / 5 / 10.1 PFLOPs) — same underlying numbers, brochure just carries one more significant digit. bf16TFLOPS is set equal to fp16TFLOPS: AMD's page lists identical PFLOP figures for \"Peak BF16 Matrix\" and \"Peak FP16 Matrix\" (both 2.5/5 PFLOPs dense/sparse), so the brochure's extra FP16 precision is assumed to carry over to BF16 as well (not separately confirmed at that precision). fp6TFLOPS/fp4TFLOPS (MXFP6/MXFP4, both 10.1 PFLOPs) and int8TOPS/int8TOPSSparse (5/10.1 POPs) come from the product page only, at its native 1-decimal-PFLOP precision — no brochure cross-check available for these.\nfp64TFLOPS and fp64TFLOPSMatrix are both 78.6 TFLOPs, and fp32TFLOPS (vector) equals the page's separately-listed \"FP32 Matrix\" figure (157.3) — AMD's page lists identical vector/matrix numbers for FP64 and FP32 on this part, unlike FP16 where vector (157.3) and matrix (2,500+) diverge sharply.\ninterconnectBandwidthGBs (153) is the page's headline \"Peak Infinity Fabric Link Bandwidth\" figure across 7 links per GPU; the page also separately lists scale-up (153 GB/s) vs. scale-out (128 GB/s, 1 link) peak figures — 153 used here as the representative headline number.\nNo msrpUSD — AMD doesn't publish Instinct list prices.\n"
    },
    {
      "id": "amd-instinct-mi455x",
      "vendor": "amd",
      "name": "Instinct MI455X",
      "category": "datacenter",
      "architecture": "CDNA 5",
      "releaseYear": 2026,
      "launchDate": "2026-07-23",
      "vramGB": 432,
      "memoryType": "HBM4",
      "memoryBandwidthGBs": 23300,
      "cacheMB": 192,
      "eccSupport": true,
      "fp64TFLOPS": 5,
      "fp32TFLOPS": 315,
      "fp16TFLOPSVector": 315,
      "fp64TFLOPSMatrix": 5,
      "bf16TFLOPS": 5000,
      "fp16TFLOPS": 5000,
      "fp8TFLOPS": 10050,
      "fp6TFLOPS": 20100,
      "fp4TFLOPS": 40300,
      "int8TOPS": 5000,
      "fp16TFLOPSSparse": 10100,
      "fp8TFLOPSSparse": 20100,
      "int8TOPSSparse": 10100,
      "processNode": "TSMC 2nm | 3nm FinFET",
      "transistorCountBillion": 320,
      "computeUnitLabel": "Work Group Processors",
      "computeUnitCount": 256,
      "peakClockMHz": 2400,
      "formFactor": "Enhanced Accelerator Module (EAM)",
      "coolingType": "Direct Liquid Cooling (DLC)",
      "interconnect": "UALink / UALoE",
      "interconnectBandwidthGBs": 3600,
      "multiInstanceSupport": "SR-IOV",
      "notes": "Part of the MI400 series / \"Helios\" rack-scale platform. AMD's product page (amd.com/en/products/accelerators/instinct/mi400/mi455x.html, \"Expand All\" spec accordion) now publishes a full spec sheet as of this pull (2026-09-01) — this collection's earlier entry, written before launch, had to derive/estimate several figures that are now directly confirmed.\nfp16TFLOPS/fp16TFLOPSSparse (5,000/10,100, \"Peak Matrix FP16 Performance\") and bf16TFLOPS (5,000, \"Peak bfloat16 (BF16) Matrix Performance\" — its 10,100 sparse variant isn't captured in a separate field) are both directly stated and identical, unlike CDNA4 where BF16 had to be inferred equal to FP16. fp8TFLOPS/fp8TFLOPSSparse are still DERIVED, not directly stated: the page lists a single \"Peak OCP FP8 Performance: 20.1 PFLOPs\" with no dense/sparse split — by analogy to the FP16/INT8 rows on this same page (both showing dense = 2x lower than their own sparse figure, and dense FP8 = 2x dense FP16 on every CDNA3/CDNA4 part in this collection), 20.1 PFLOPs is treated as the SPARSE figure, giving a derived dense fp8TFLOPS of ~10,050 — same estimate this file used pre-launch, now more confidently cross-checked against the newly-published FP16/INT8 dense:sparse ratios. fp6TFLOPS (MXFP6, 20,100) and fp4TFLOPS (MXFP4, 40,300) are used as directly published, single-figure, precision unconfirmed as dense or sparse.\nfp64 is surprisingly low on this part — both vector and matrix rated at just 5 TFLOPs (vs. 78.6 on MI355X) — taken directly from the page as published, not adjusted; MI455X is evidently not built for FP64/HPC workloads the way CDNA3/4 parts were. fp32 vector and matrix are both 315 TFLOPs (also equal to the page's separately-stated FP16 vector figure).\ncomputeUnitLabel is \"Work Group Processors\" (256) — AMD's new CDNA5 terminology replacing \"Compute Units\"; no separate Stream Processor/Matrix Core counts are published for this part. interconnectBandwidthGBs (3,600) uses the page's \"Scale-up (Peak) UALoE Bi-directional Bandwidth\" figure as the headline number; the complementary \"Scale-out (Peak) UALink Bi-directional Bandwidth\" (600 GB/s) is a separate, smaller network-fabric figure not captured in its own field.\ntdpWatts and targetId are still both omitted: no TBP or PCIe bus type is published on this page (unofficial estimates put TBP \"north of 2kW\" based on Helios rack power budgets — not used here), and ROCm/CDNA5 support isn't reflected in this site's `rocm` collection yet, so there's no confirmed gfx target to show. No msrpUSD — sold as part of rack-scale deals, not a per-unit list price.\n"
    },
    {
      "id": "amd-radeon-ai-pro-r9700",
      "vendor": "amd",
      "name": "Radeon AI PRO R9700",
      "category": "workstation",
      "architecture": "RDNA 4",
      "releaseYear": 2025,
      "targetId": "gfx1201",
      "vramGB": 32,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 640,
      "cacheMB": 64,
      "eccSupport": true,
      "fp32TFLOPS": 47.8,
      "fp16TFLOPSVector": 47.8,
      "fp16TFLOPS": 191,
      "fp8TFLOPS": 383,
      "int8TOPS": 383,
      "int4TOPS": 766,
      "fp16TFLOPSSparse": 383,
      "fp8TFLOPSSparse": 766,
      "int8TOPSSparse": 766,
      "transistorCountBillion": 53.9,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 64,
      "shaderCoreCount": 4096,
      "matrixCoreCount": 128,
      "peakClockMHz": 2920,
      "tdpWatts": 300,
      "formFactor": "Double-slot desktop add-in card",
      "coolingType": "Active",
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "PCIe only (no dedicated multi-GPU link)",
      "notes": "The largest-VRAM AMD card outside the Instinct line, and the reason it matters here: 32 GB on gfx1201, the same Navi 48 target as the RX 9070 XT, so it inherits that target's ROCm support from 6.4.1 (2025-05) onward. Anyone hitting the 16 GB ceiling on RDNA 4 consumer cards gets double the capacity without changing gfx target, ROCm version, or toolchain — see /rocm-consumer-gpus/ for the consumer-side picture.\nAdded 2026-09-21 after a reader pointed out the site covered NVIDIA's RTX PRO 6000 Blackwell but had no AMD workstation entry at all — a vendor-coverage asymmetry, not an oversight about one card.\nCompute specs are identical to the RX 9070 XT's full Navi 48 configuration (64 CUs, 4096 stream processors, 128 AI accelerators, 53.9 B transistors) at slightly lower clocks (2920 MHz boost vs 2970) and 300 W instead of 304 W. What actually differs is memory capacity (32 GB vs 16 GB on the same 256-bit bus at the same 640 GB/s) and ECC support, which AMD publishes as \"Yes (Linux Only)\" — recorded as `eccSupport: true` with that OS restriction noted here rather than in a field, since the schema has no place for it. Bandwidth being identical to the 9070 XT's means this card fits much larger models without feeding them any faster.\nfp16TFLOPS (191, dense) and fp8TFLOPS (383, dense) are directly published, each with an explicit \"with Structured Sparsity\" variant (383 / 766) in the ...Sparse fields — AMD publishes dense as the headline and sparsity separately, the opposite of NVIDIA's convention, so no halving is applied. int8TOPS/int8TOPSSparse mirror the FP8 figures (383/766); int4TOPS is listed but its sparsity variant (1,531 TOPs) has no schema field. matrixCoreCount uses the page's \"AI Accelerators\" count (128).\nAMD also sells a spec-identical R9700S at amd-radeon-ai-pro-r9700s.html — every compute, memory and power figure matches this card exactly; the only published difference is Cooling (Passive vs Active), i.e. an OEM/SI variant for chassis with their own airflow. Deliberately not given its own entry: it shares this card's gfx target, capacity and throughput, so it would add a row without adding a compatibility fact. (Contrast the RX 9060 XT 8GB/16GB pair, which do get separate entries, because VRAM capacity changes which models fit.)\npeakClockMHz uses the published Boost Frequency (\"Up to 2920 MHz\"), not the 2350 MHz Game Frequency. tdpWatts is the published \"Total Board Power (TBP)\". No Launch Date field is published on this page, so launchDate is omitted; releaseYear (2025) is unaffected — AMD announced the card at COMPUTEX 2025 (2025-05-20) without naming an availability date. AMD publishes no MSRP on the spec page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html (\"Expand All\" spec accordion), pulled 2026-09-21.\n"
    },
    {
      "id": "amd-radeon-instinct-mi25",
      "vendor": "amd",
      "name": "Radeon Instinct MI25",
      "category": "datacenter",
      "architecture": "Vega10 (GCN 5.0)",
      "releaseYear": 2017,
      "launchDate": "2017 (AMD's own performance footnote dated Jun 2, 2017)",
      "targetId": "gfx900",
      "vramGB": 16,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 484,
      "eccSupport": true,
      "fp64TFLOPS": 0.768,
      "fp32TFLOPS": 12.3,
      "fp16TFLOPSVector": 24.6,
      "processNode": "14nm FinFET",
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 64,
      "shaderCoreCount": 4096,
      "tdpWatts": 300,
      "formFactor": "Full-Height, Dual-Slot",
      "coolingType": "Passive",
      "pcieGen": "PCIe Gen3 x16",
      "multiInstanceSupport": "MxGPU SR-IOV (hardware virtualization, not MIG-style GPU partitioning)",
      "notes": "Added 2026-09 as AMD's first Radeon Instinct card (predates this collection's previous oldest AMD entry, 2020's MI100, by three years) — the user specifically asked for older Radeon Instinct cards alongside NVIDIA's K80/M40. AMD's own current site no longer hosts this datasheet (amd.com/system/files/documents/radeon-instinct-mi25-datasheet.pdf 404s); pulled via the Wayback Machine's 2018 archive of that exact URL, read via pdftotext — an archived copy of AMD's own document, not a third-party remake.\nfp16TFLOPSVector (24.6) and fp32TFLOPS (12.3) are Vega10's \"Rapid Packed Math\" throughput via its standard stream processors — correctly placed in this schema's vector field, not fp16TFLOPS/matrix, because Vega predates AMD's dedicated Matrix Cores entirely (introduced with CDNA/ MI100 three years later). No INT8 figure is published for this card at all (unlike the later MI50) — matrix/tensor-throughput fields are correctly all omitted, not zero. fp64TFLOPS (0.768, i.e. \"768 GFLOPS\") is explicitly stated as a 1/16th-of-FP32 rate in AMD's own footnote. eccSupport is \"Yes\" but per AMD's own footnote 3, \"ECC support is limited to the HBM2 memory and ECC protection is not provided for internal GPU structures\" — partial, unlike MI50's full-chip ECC the following generation.\nNo Infinity Fabric / GPU-to-GPU interconnect field populated — Infinity Fabric Link was introduced with MI50/MI60 the following generation; MI25 is PCIe Gen3 x16 only for both host and (if used) peer-to-peer traffic, confirmed by MI50's own datasheet footnote describing \"previous Gen Radeon Instinct\" cards as PCIe-Gen3-only. targetId (gfx900) doesn't cross-link to any `rocm` entry in this collection — every ROCm version tracked here (oldest: 5.0.0) already dropped Vega10/gfx900 support, so the empty ROCm-compatibility section on this GPU's page is accurate, not a data gap (the same situation as this round's K80 addition on the CUDA side). No transistor count or MSRP — AMD doesn't publish Instinct list prices (see this collection's MI100 entry for the same note).\n"
    },
    {
      "id": "amd-radeon-instinct-mi50",
      "vendor": "amd",
      "name": "Radeon Instinct MI50",
      "category": "datacenter",
      "architecture": "Vega 7nm (GCN 5.1)",
      "releaseYear": 2018,
      "launchDate": "Nov 2018 (announced); datasheet copyright 2019, performance footnote dated Jun 19, 2019",
      "targetId": "gfx906",
      "vramGB": 32,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 1000,
      "eccSupport": true,
      "fp64TFLOPS": 6.6,
      "fp32TFLOPS": 13.3,
      "fp16TFLOPSVector": 26.5,
      "int8TOPS": 53.6,
      "processNode": "7nm FinFET",
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 60,
      "shaderCoreCount": 3840,
      "tdpWatts": 300,
      "formFactor": "Full-Height, Dual-Slot",
      "coolingType": "Passive",
      "pcieGen": "PCIe 3.0/4.0 x16 (Gen4 capable)",
      "interconnect": "Infinity Fabric Link (dual)",
      "interconnectBandwidthGBs": 184,
      "multiInstanceSupport": "MxGPU SR-IOV (hardware virtualization, not MIG-style GPU partitioning)",
      "notes": "Added 2026-09 — this collection's own README had specifically flagged MI50 as a candidate older Radeon Instinct card, and its gfx906 target is already referenced across roughly twenty `rocm` entries in this collection (4.5.x through 6.3.3, deprecated from 6.0.0) without a corresponding `gpus` entry until now. All figures sourced directly from AMD's official \"AMD Radeon Instinct MI50 (16GB & 32GB)\" datasheet (amd.com/system/files/documents/ radeon-instinct-mi50-datasheet.pdf), read via pdftotext from a mirrored copy after amd.com's own URL stopped serving the file directly (same situation as several AMD product pages elsewhere in this collection); content matches AMD's own footnote dating (Jun 19, 2019) and copyright year, so treated as a faithful copy of AMD's original.\nThis entry represents the 32GB variant — the datasheet states \"16GB or 32GB HBM2\" as a single row with identical compute/bandwidth/power specs for both capacities, so vramGB (32) is the only field that would differ; matches this site's general preference for the higher-memory SKU where one datasheet covers both (see V100 SXM2 32GB, M40 24GB). fp16TFLOPSVector (26.5) and fp32TFLOPS (13.3) are Vega 7nm's Rapid-Packed-Math throughput via standard stream processors, correctly in this schema's vector field — Vega still predates AMD's dedicated Matrix Cores (introduced with CDNA/MI100 the following generation). int8TOPS (53.6) is likewise a CU-based packed-math rate, not Matrix-Core throughput, but is placed in this schema's matrix/tensor-throughput field since no vector-INT8 field exists — noted here for anyone comparing it against later CDNA entries' genuinely Matrix-Core-derived INT8 figures. eccSupport is \"Yes\" and, per the datasheet's own \"ECC (Full-chip): Yes\" row, full-chip (HBM2 plus internal GPU structures) — an upgrade over MI25's HBM2-only ECC the prior generation.\ninterconnectBandwidthGBs (184) is the dual Infinity-Fabric-Link GPU-to-GPU bandwidth specifically (AMD's footnote: \"184 GB/s peak theoretical GPU to GPU ... per GPU card\"), distinct from pcieGen's host link; AMD's footnote separately states these combine to a 248 GB/s aggregate I/O figure when both are counted together, not modeled as a separate field here. No transistor count or MSRP — AMD doesn't publish Instinct list prices (see this collection's MI100/MI25 entries).\n"
    },
    {
      "id": "amd-radeon-instinct-mi6",
      "vendor": "amd",
      "name": "Radeon Instinct MI6",
      "category": "datacenter",
      "architecture": "Polaris (GCN 4)",
      "releaseYear": 2017,
      "launchDate": "June 2017",
      "targetId": "gfx803",
      "vramGB": 16,
      "memoryType": "GDDR5",
      "memoryBandwidthGBs": 224,
      "eccSupport": false,
      "fp64TFLOPS": 0.358,
      "fp32TFLOPS": 5.73,
      "fp16TFLOPSVector": 5.73,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 36,
      "shaderCoreCount": 2304,
      "tdpWatts": 150,
      "formFactor": "Full-Height, Single-Slot",
      "coolingType": "Passive",
      "pcieGen": "PCIe 3.0 x16",
      "notes": "Added 2026-09 as the entry-level member of AMD's original three-card Radeon Instinct launch lineup (alongside MI8 and this collection's existing MI25), predating this collection's other Radeon Instinct entries by nothing in calendar time (all three launched June 2017) but representing an older, lower-tier architecture — Polaris (GCN 4th Gen) rather than MI25's Vega10 (GCN 5.0). All figures read directly from AMD's own live, current \"Accelerator Specifications\" comparison tool (amd.com/en/products/specifications/accelerators.html), which still carries structured spec data for this decade-old card; extracted from the page's embedded JSON data model (not every field AMD's schema defines is populated for this old a product — e.g. Compute Units and Matrix Cores are present as columns but empty for MI6, so neither was taken from that source — Compute Units was later derived from the published shader-core count, see below, and Matrix Cores stays empty).\nfp16TFLOPSVector (5.73) is numerically identical to fp32TFLOPS (5.73) in AMD's own table — Polaris predates Rapid Packed Math (introduced with Vega/GCN 5 the following generation), so there's no 2x FP16 speedup and no separate matrix/tensor throughput field applies (Polaris has no Matrix Cores at all, introduced with CDNA/MI100 years later). eccSupport is \"no\" per AMD's own table (Memory ECC Support: No) — this cost-optimized entry-tier card uses plain GDDR5 without the ECC that MI8's HBM and MI25/MI50's HBM2 carry. targetId (gfx803) is shared with MI8 (both Polaris10 and Fiji use the same gfx803 LLVM/ROCm target ISA, per LLVM's AMDGPU backend docs) but doesn't cross-link to any `rocm` entry in this collection — every ROCm version tracked here (oldest: 5.0.0) already dropped gfx803 support, so the empty ROCm-compatibility section is accurate, matching this round's K80/MI25 precedent. No MxGPU/multi-instance field — not listed for this generation (introduced starting with MI25/MI50). No transistor count, process node, or MSRP — AMD doesn't publish Instinct list prices (see this collection's other Instinct entries).\n2026-09-15 — added computeUnitCount (36 Compute Units). AMD's own specifications tool leaves its Compute Units column empty for this card (noted above), but 36 is not a guess: GCN fixes a Compute Unit at 64 shader cores, so the 2,304 shaderCoreCount AMD does publish gives 2,304 / 64 = 36. It is the same CU x 64 relationship this collection's MI210 entry already documents, applied in the opposite direction. matrixCoreCount stays empty — Polaris has no Matrix Cores at all.\n"
    },
    {
      "id": "amd-radeon-instinct-mi60",
      "vendor": "amd",
      "name": "Radeon Instinct MI60",
      "category": "datacenter",
      "architecture": "Vega 7nm (GCN 5.1)",
      "releaseYear": 2018,
      "launchDate": "Nov 6, 2018 (announced; AMD said shipping to datacenter customers by end of 2018)",
      "targetId": "gfx906",
      "vramGB": 32,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 1000,
      "eccSupport": true,
      "fp64TFLOPS": 7.4,
      "fp32TFLOPS": 14.8,
      "fp16TFLOPSVector": 29.5,
      "processNode": "7nm FinFET",
      "transistorCountBillion": 13.2,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 64,
      "shaderCoreCount": 4096,
      "tdpWatts": 300,
      "pcieGen": "PCIe 4.0 x16",
      "interconnect": "Infinity Fabric Link (dual)",
      "interconnectBandwidthGBs": 200,
      "multiInstanceSupport": "MxGPU SR-IOV (hardware virtualization, not MIG-style GPU partitioning)",
      "notes": "Added 2026-09 — the full-die sibling of the MI50 (same Vega 20 / gfx906 silicon), and a card that comes up constantly in local-LLM discussions alongside it. Shares the MI50's ROCm history exactly, since ROCm grants support per gfx target: see the MI50 and Radeon VII entries.\nFigures sourced from AMD's official launch press release of 2018-11-06 (\"AMD Unveils World's First 7nm Datacenter GPUs\", as distributed via GlobeNewswire): 29.5 TFLOPS FP16, 14.8 TFLOPS FP32, 7.4 TFLOPS FP64 peak theoretical, 13.2 billion transistors on a 331.46mm² die at 300W (footnote 1); 32GB HBM2 with full-chip ECC covering \"HBM2 memory and internal GPU structures\" (footnote 6); dual Infinity Fabric Links at \"up to 200 GB/s peak theoretical GPU to GPU\" per card, 264 GB/s aggregate with PCIe Gen 4 (footnote 4); MxGPU hardware virtualization. memoryBandwidthGBs (1000) follows the release's \"up to 1 TB/s\" wording, matching how the MI50 entry records the same memory subsystem. AMD's MI60 datasheet PDF (amd.com/system/files/documents/ radeon-instinct-mi60-datasheet.pdf) could not be retrieved at time of writing.\ncomputeUnitCount (64) and shaderCoreCount (4096) are not stated in the press release: they are the full Vega 20 die (64 stream processors per CU; the MI50's datasheet lists 60 CUs / 3840 for the cut-down part), and are consistent with AMD's own FP32 figure — 4096 x 2 FLOP x 1.8 GHz = 14.7 TFLOPS. No INT8 figure is given in the press release, so int8TOPS is omitted rather than inferred. fp16TFLOPSVector is Vega's Rapid-Packed-Math rate on the stream processors: gfx906 predates AMD's Matrix Cores (introduced with CDNA/MI100).\n"
    },
    {
      "id": "amd-radeon-instinct-mi8",
      "vendor": "amd",
      "name": "Radeon Instinct MI8",
      "category": "datacenter",
      "architecture": "Fiji (GCN 3)",
      "releaseYear": 2017,
      "launchDate": "June 2017",
      "targetId": "gfx803",
      "vramGB": 4,
      "memoryType": "HBM",
      "memoryBandwidthGBs": 512,
      "eccSupport": false,
      "fp64TFLOPS": 0.512,
      "fp32TFLOPS": 8.9,
      "fp16TFLOPSVector": 8.9,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 64,
      "shaderCoreCount": 4096,
      "tdpWatts": 175,
      "formFactor": "Full-Height, Dual-Slot",
      "coolingType": "Passive",
      "pcieGen": "PCIe 3.0 x16",
      "notes": "Added 2026-09 alongside MI6 — the mid-tier card in AMD's original three-card Radeon Instinct launch lineup (with MI6 entry-level and this collection's existing MI25 as flagship, all launched June 2017). Built on Fiji (same die as the Radeon R9 Fury X/Nano), AMD's first HBM-equipped GPU architecture, hence the unusually small 4GB VRAM ceiling for a \"datacenter\" card — first-generation HBM's stacks topped out at 1GB each with only 4 stacks used. All figures read directly from AMD's own live \"Accelerator Specifications\" comparison tool (amd.com/en/products/specifications/accelerators.html), extracted from the page's embedded JSON data model, 2026-09 — the same source as MI6's entry. As with MI6, AMD's own schema defines Compute Units and Matrix Cores columns that are simply empty for this old a product, so neither figure was taken from that source; Compute Units was later derived from the published shader-core count (see below) and Matrix Cores stays empty.\nfp16TFLOPSVector (8.9) is numerically identical to fp32TFLOPS (8.9) in AMD's own table — Fiji (GCN 3rd Gen) predates Rapid Packed Math (introduced with Vega/GCN 5) even further back than MI6's Polaris, so there's no 2x FP16 speedup here either, and no matrix/tensor field applies (Fiji has no Matrix Cores, introduced with CDNA/MI100 years later). eccSupport is \"no\" per AMD's own table — unlike MI25/MI50's HBM2 the following generation, this first-gen HBM implementation carries no ECC. targetId (gfx803) is shared with MI6 (Fiji and Polaris10 use the same gfx803 LLVM/ROCm target ISA) but doesn't cross-link to any `rocm` entry in this collection — every ROCm version tracked here already dropped gfx803 support, matching this round's MI6/K80/MI25 precedent. No MxGPU/multi-instance field, transistor count, process node, or MSRP — see MI6's entry for the same reasoning.\n2026-09-15 — added computeUnitCount (64 Compute Units), derived exactly as MI6's was: GCN fixes a Compute Unit at 64 shader cores, so the 4,096 shaderCoreCount from AMD's own table gives 4,096 / 64 = 64 — the full Fiji die with nothing fused off. matrixCoreCount stays empty; Fiji has no Matrix Cores.\n"
    },
    {
      "id": "amd-radeon-rx-6600",
      "vendor": "amd",
      "name": "Radeon RX 6600",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2021,
      "launchDate": "2021-10-13",
      "targetId": "gfx1032",
      "vramGB": 8,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 224,
      "cacheMB": 32,
      "fp32TFLOPS": 8.93,
      "fp16TFLOPSVector": 17.86,
      "transistorCountBillion": 11.1,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 28,
      "shaderCoreCount": 1792,
      "peakClockMHz": 2491,
      "tdpWatts": 132,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Deliberately included as an *unsupported* card, same reasoning as the RX 6700 XT. This is Navi 23 = gfx1032, which appears in no ROCm release's supported-target list in this collection. It is a very widely owned budget card, so \"does ROCm work on my 6600?\" is a common question with a firm answer: not officially, ever. The community route is HSA_OVERRIDE_GFX_VERSION=10.3.0 (borrowing gfx1030's ABI from the same RDNA 2 family) — unofficial, self-managed, and not reflected in this site's compatibility data. Note also that 8 GB of VRAM is the binding constraint for local LLM work regardless of ROCm status; run it through /vram-calculator/ before planning around it.\nfp16TFLOPS and fp8TFLOPS intentionally omitted — RDNA 2 has no matrix/AI-accelerator cores; AMD's page lists only \"Peak Half Precision (FP16 Vector) Performance: 17.86 TFLOPs\" (captured as fp16TFLOPSVector) with no \"AI Accelerators\" count, so matrixCoreCount is omitted too.\npeakClockMHz uses the published Boost Frequency (\"Up to 2491 MHz\"), not the 2044 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages, so pcieGen is left unset (this card is commonly reported as a PCIe 4.0 x8 part, but AMD does not state a bus width on the spec page, so nothing is recorded here). AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6600.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-6650-xt",
      "vendor": "amd",
      "name": "Radeon RX 6650 XT",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2022,
      "launchDate": "2022-05-10",
      "targetId": "gfx1032",
      "vramGB": 8,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 280,
      "cacheMB": 32,
      "fp32TFLOPS": 10.79,
      "fp16TFLOPSVector": 21.59,
      "transistorCountBillion": 11.1,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 32,
      "shaderCoreCount": 2048,
      "peakClockMHz": 2635,
      "tdpWatts": 180,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Included as an unsupported card, same as the RX 6600 it shares a die with. Navi 23 = gfx1032, which appears in no ROCm release's supported-target list. Added 2026-09-21 alongside the RX 6750 XT to complete the two unsupported RDNA 2 target families, so someone searching a specific SKU finds a page rather than a gap.\nThe community route is HSA_OVERRIDE_GFX_VERSION=10.3.0, borrowing gfx1030's ABI from the same RDNA 2 family — unofficial, self-managed, and not reflected in this site's compatibility data. Note that 8 GB of VRAM binds before ROCm status does for local LLM work; run anything you plan to load through /vram-calculator/ first.\nA refresh of the RX 6600 XT rather than a new design: same 32 CUs, 2048 stream processors, 11.1 B transistors and 32 MB Infinity Cache, with faster 17.5 Gbps memory (280 GB/s) and a 180 W board power.\nfp16TFLOPS and fp8TFLOPS intentionally omitted — RDNA 2 has no matrix/AI-accelerator cores; AMD's page lists only \"Peak Half Precision (FP16 Vector) Performance: 21.59 TFLOPs\" (packed-math vector rate, captured as fp16TFLOPSVector) with no \"AI Accelerators\" count, so matrixCoreCount is omitted too.\npeakClockMHz uses the published Boost Frequency (\"Up to 2635 MHz\"), not the 2410 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages, so pcieGen is left unset. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6650-xt.html (\"Expand All\" spec accordion), pulled 2026-09-21.\n"
    },
    {
      "id": "amd-radeon-rx-6700-xt",
      "vendor": "amd",
      "name": "Radeon RX 6700 XT",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2021,
      "launchDate": "2021-03-01",
      "targetId": "gfx1031",
      "vramGB": 12,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 384,
      "cacheMB": 96,
      "fp32TFLOPS": 13.21,
      "fp16TFLOPSVector": 26.43,
      "transistorCountBillion": 17.2,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 40,
      "shaderCoreCount": 2560,
      "peakClockMHz": 2581,
      "tdpWatts": 230,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Deliberately included as an *unsupported* card. This is Navi 22 = gfx1031, which appears in no ROCm release's supported-target list — not in any 5.x, 6.x, 7.x or 10.x entry in this collection. That is not a data gap: ROCm grants support per gfx target, AMD has never validated gfx1031, and the 6700 XT / 6750 XT are among the most common cards people try to run ROCm on anyway, which is exactly why the HSA_OVERRIDE_GFX_VERSION workaround exists (typically HSA_OVERRIDE_GFX_VERSION=10.3.0, borrowing the gfx1030 ABI from the same RDNA 2 family). Unofficial and self-managed — see /rocm-consumer-gpus/ for the caveats.\nfp16TFLOPS and fp8TFLOPS intentionally omitted — RDNA 2 has no matrix/AI-accelerator cores; AMD's page lists only \"Peak Half Precision (FP16 Vector) Performance: 26.43 TFLOPs\" (captured as fp16TFLOPSVector) with no \"AI Accelerators\" count, so matrixCoreCount is omitted too.\npeakClockMHz uses the published Boost Frequency (\"Up to 2581 MHz\"), not the 2424 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages, so pcieGen is left unset. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6700-xt.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-6750-xt",
      "vendor": "amd",
      "name": "Radeon RX 6750 XT",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2022,
      "launchDate": "2022-05-10",
      "targetId": "gfx1031",
      "vramGB": 12,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 432,
      "cacheMB": 96,
      "fp32TFLOPS": 13.31,
      "fp16TFLOPSVector": 26.62,
      "transistorCountBillion": 17.2,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 40,
      "shaderCoreCount": 2560,
      "peakClockMHz": 2600,
      "tdpWatts": 250,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Included as an unsupported card, same as the RX 6700 XT it shares a die with. Navi 22 = gfx1031, which appears in no ROCm release's supported-target list.\nAdded 2026-09-21 in direct response to an r/ROCm commenter who runs ROCm on this exact card via the override trick and found it missing from the site — the 6700 XT was listed but this refresh of the same chip was not. Their report (that it works when made to present as an RX 6800) is consistent with what /rocm-consumer-gpus/ describes: HSA_OVERRIDE_GFX_VERSION=10.3.0 borrows gfx1030's ABI from the same RDNA 2 family. That remains an unofficial, self-managed workaround and is deliberately not reflected in this site's compatibility data, which tracks what AMD validates.\nA refresh of the 6700 XT rather than a new design: same 40 CUs, 2560 stream processors, 17.2 B transistors and 96 MB Infinity Cache, with faster 18 Gbps memory (432 GB/s vs 384), higher clocks (2600 MHz vs 2581 boost) and a higher board power (250 W vs 230 W).\nfp16TFLOPS and fp8TFLOPS intentionally omitted — RDNA 2 has no matrix/AI-accelerator cores; AMD's page lists only \"Peak Half Precision (FP16 Vector) Performance: 26.62 TFLOPs\" (packed-math vector rate, captured as fp16TFLOPSVector) with no \"AI Accelerators\" count, so matrixCoreCount is omitted too.\npeakClockMHz uses the published Boost Frequency (\"Up to 2600 MHz\"), not the 2495 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages, so pcieGen is left unset. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6750-xt.html (\"Expand All\" spec accordion), pulled 2026-09-21.\n"
    },
    {
      "id": "amd-radeon-rx-6800",
      "vendor": "amd",
      "name": "Radeon RX 6800",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2020,
      "launchDate": "2020-11-18",
      "targetId": "gfx1030",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 512,
      "cacheMB": 128,
      "fp32TFLOPS": 16.17,
      "fp16TFLOPSVector": 32.33,
      "transistorCountBillion": 26.8,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 60,
      "shaderCoreCount": 3840,
      "peakClockMHz": 2105,
      "tdpWatts": 250,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Shares the Navi 21 die (and therefore gfx1030) with the RX 6800 XT / 6900 XT / 6950 XT, so ROCm's gfx1030 support covers this card too even though AMD's compatibility matrix never names it — the same reason gpuNames.ts describes itself as flagship names only, not an exhaustive SKU list.\nfp16TFLOPS and fp8TFLOPS intentionally omitted, not merely unconfirmed: RDNA 2 has no dedicated matrix/AI-accelerator cores (introduced in RDNA 3), so there is no matrix-throughput figure comparable to the rest of this collection. AMD's page confirms by omission — it lists only \"Peak Half Precision (FP16 Vector) Performance: 32.33 TFLOPs\" (packed-math vector rate, 2x FP32, captured as fp16TFLOPSVector) with no \"FP16 Matrix\" row, and no \"AI Accelerators\" count, so matrixCoreCount is omitted as well. Same treatment as the RX 6900 XT entry.\npeakClockMHz uses the published Boost Frequency (\"Up to 2105 MHz\"), not the 1815 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages (unlike 9000-series), so pcieGen is left unset rather than assumed. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6800.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-6800-xt",
      "vendor": "amd",
      "name": "Radeon RX 6800 XT",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2020,
      "launchDate": "2020-11-18",
      "targetId": "gfx1030",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 512,
      "cacheMB": 128,
      "fp32TFLOPS": 20.74,
      "fp16TFLOPSVector": 41.47,
      "transistorCountBillion": 26.8,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 72,
      "shaderCoreCount": 4608,
      "peakClockMHz": 2250,
      "tdpWatts": 300,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Same Navi 21 die (gfx1030) as the RX 6800 / 6900 XT / 6950 XT, so it inherits ROCm's gfx1030 support without being named in AMD's compatibility matrix.\nfp16TFLOPS and fp8TFLOPS intentionally omitted — RDNA 2 has no matrix/AI-accelerator cores, so there is no matrix figure comparable to the rest of this collection; AMD's page lists only \"Peak Half Precision (FP16 Vector) Performance: 41.47 TFLOPs\" (captured as fp16TFLOPSVector) and publishes no \"AI Accelerators\" count, so matrixCoreCount is omitted too. Same treatment as the RX 6900 XT entry.\npeakClockMHz uses the published Boost Frequency (\"Up to 2250 MHz\"), not the 2015 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages, so pcieGen is left unset. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6800-xt.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-6900-xt",
      "vendor": "amd",
      "name": "Radeon RX 6900 XT",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2020,
      "launchDate": "2020-12-08",
      "targetId": "gfx1030",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 512,
      "cacheMB": 128,
      "fp32TFLOPS": 23.04,
      "fp16TFLOPSVector": 46.08,
      "transistorCountBillion": 26.8,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 80,
      "shaderCoreCount": 5120,
      "peakClockMHz": 2250,
      "tdpWatts": 300,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "fp16TFLOPS and fp8TFLOPS are both intentionally omitted, not just unconfirmed — RDNA 2 has no dedicated matrix/AI-accelerator cores at all (those were introduced with RDNA 3), so there's no matrix-throughput figure to report that's methodologically comparable to every other entry in this collection. AMD's own product page confirms this by omission: it lists only \"Peak Half Precision (FP16 Vector) Performance: 46.08 TFLOPs\" (packed-math vector rate, 2x FP32, not a matrix/tensor figure, captured here as fp16TFLOPSVector) with no \"FP16 Matrix\" row at all, unlike every RDNA3/RDNA4 card in this collection. Included anyway (unlike, say, a card with zero ROCm support) because ROCm officially supports gfx1030 and this site's whole point is covering the ROCm side other trackers skip — VRAM/bandwidth/TDP are still useful even without a compute-throughput number.\nNo pcieGen or eccSupport field — AMD's consumer Radeon spec page doesn't publish a PCIe bus generation or ECC-support line the way Instinct pages do (only \"Additional Power Connector: 2x8-Pin\" under General); left unconfirmed rather than assumed. formFactor similarly omitted — the page's \"Board Type: Component\" isn't a meaningful form-factor label like Instinct's \"OAM Module\"/\"PCIe Add-in Card\". Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6900-xt.html (\"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-radeon-rx-6950-xt",
      "vendor": "amd",
      "name": "Radeon RX 6950 XT",
      "category": "consumer",
      "architecture": "RDNA 2",
      "releaseYear": 2022,
      "launchDate": "2022-05-10",
      "targetId": "gfx1030",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 576,
      "cacheMB": 128,
      "fp32TFLOPS": 23.65,
      "fp16TFLOPSVector": 47.31,
      "transistorCountBillion": 26.8,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 80,
      "shaderCoreCount": 5120,
      "peakClockMHz": 2310,
      "tdpWatts": 335,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "The fastest gfx1030 card — same Navi 21 die and CU count as the RX 6900 XT, with higher clocks (2310 MHz vs 2250 MHz boost) and faster 18 Gbps memory (576 GB/s vs 512 GB/s). Inherits ROCm's gfx1030 support.\nfp16TFLOPS and fp8TFLOPS intentionally omitted — RDNA 2 has no matrix/AI-accelerator cores; AMD's page lists only \"Peak Half Precision (FP16 Vector) Performance: 47.31 TFLOPs\" (captured as fp16TFLOPSVector) with no \"FP16 Matrix\" row and no \"AI Accelerators\" count, so matrixCoreCount is omitted too. Same treatment as the RX 6900 XT entry.\npeakClockMHz uses the published Boost Frequency (\"Up to 2310 MHz\"), not the 2100 MHz Game Frequency. No PCIe \"Bus Type\" line is published on 6000-series pages, so pcieGen is left unset. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/6000-series/amd-radeon-rx-6950-xt.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-7600-xt",
      "vendor": "amd",
      "name": "Radeon RX 7600 XT",
      "category": "consumer",
      "architecture": "RDNA 3",
      "releaseYear": 2024,
      "launchDate": "2024-01-24",
      "targetId": "gfx1102",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 288,
      "cacheMB": 32,
      "fp32TFLOPS": 22.6,
      "fp16TFLOPSVector": 22.6,
      "fp16TFLOPS": 45.1,
      "int8TOPS": 45.1,
      "int4TOPS": 90.2,
      "transistorCountBillion": 13.3,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 32,
      "shaderCoreCount": 2048,
      "matrixCoreCount": 64,
      "peakClockMHz": 2755,
      "tdpWatts": 190,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Navi 33 = gfx1102, and the most interesting ROCm timeline of any card in this collection: gfx1102 appears in no ROCm release before 7.14.0, which means this January-2024 card waited until 2026 for official support — the longest launch-to-support gap of any consumer Radeon tracked here, and the exact opposite of the shrinking-gap trend the RDNA 4 cards show. Check the /rocm/ version list before assuming an older ROCm install covers it.\nNotable for local LLM work as the cheapest 16 GB Radeon: it pairs 7600-class compute with double the VRAM, so it fits models that the faster 12 GB 7700 XT cannot. The tradeoff is bandwidth — 288 GB/s on a 128-bit bus, the lowest of any 16 GB card here, which bounds token generation speed even when a model fits.\nfp16TFLOPS (45.1, dense) is AMD's published \"Peak Half Precision (FP16 Matrix) Performance\"; the 22.6 TFLOPs vector rate is captured separately as fp16TFLOPSVector. No structured-sparsity variants (RDNA 3 has no 2:4 sparsity). fp8TFLOPS omitted — no native FP8 matrix support on RDNA 3; INT8/INT4 matrix rates published instead via the card's 64 \"AI Accelerators\" (used as matrixCoreCount).\npeakClockMHz uses the published Boost Frequency (\"Up to 2755 MHz\"), not the 2470 MHz Game Frequency. AMD's product page publishes no Launch Date field; launchDate (2024-01-24) comes from AMD's own announcement press release of 2024-01-08, which states availability beginning January 24, 2024. That release prices the card \"under $350\" without naming a figure on the spec page, so msrpUSD is omitted rather than filled from secondary reporting. Source: amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7600-xt.html (\"Expand All\" spec accordion) plus amd.com/en/newsroom/press-releases/2024-1-8-amd-unveils-amd-radeon-rx-7600-xt-graphics-card--.html for the date, pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-7700-xt",
      "vendor": "amd",
      "name": "Radeon RX 7700 XT",
      "category": "consumer",
      "architecture": "RDNA 3",
      "releaseYear": 2023,
      "launchDate": "2023-09-06",
      "targetId": "gfx1101",
      "vramGB": 12,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 432,
      "cacheMB": 48,
      "fp32TFLOPS": 35.2,
      "fp16TFLOPSVector": 35.2,
      "fp16TFLOPS": 70.3,
      "int8TOPS": 70.3,
      "int4TOPS": 141,
      "transistorCountBillion": 28.1,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 54,
      "shaderCoreCount": 3456,
      "matrixCoreCount": 108,
      "peakClockMHz": 2544,
      "tdpWatts": 245,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Same Navi 32 die (gfx1101) as the RX 7800 XT, so it inherits that card's ROCm support from 6.3.0 onward — a year after its own launch, and three releases after the gfx1100 cards it shares an architecture generation with. A good illustration of why this site tracks gfx targets rather than architecture names: \"RDNA 3 is supported\" is not a usable answer.\nfp16TFLOPS (70.3, dense) is AMD's published \"Peak Half Precision (FP16 Matrix) Performance\"; the 35.2 TFLOPs vector rate is captured separately as fp16TFLOPSVector. No structured-sparsity variants (RDNA 3 has no 2:4 sparsity). fp8TFLOPS omitted — no native FP8 matrix support on RDNA 3; INT8/INT4 matrix rates published instead via the card's 108 \"AI Accelerators\" (used as matrixCoreCount).\nmemoryBandwidthGBs uses the raw GDDR6 figure (432 GB/s), not an Infinity Cache-effective number. peakClockMHz uses the published Boost Frequency (\"Up to 2544 MHz\"), not the 2171 MHz Game Frequency. AMD's product page publishes no Launch Date field for this card; launchDate (2023-09-06) comes from AMD's own Gamescom press release of 2023-08-25, which states availability beginning September 6, 2023. AMD publishes no MSRP on the spec page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700-xt.html (\"Expand All\" spec accordion) plus amd.com/en/newsroom/press-releases/2023-8-25-new-amd-radeon-rx-7800-xt-and-radeon-rx-7700-xt-gr.html for the date, pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-7800-xt",
      "vendor": "amd",
      "name": "Radeon RX 7800 XT",
      "category": "consumer",
      "architecture": "RDNA 3",
      "releaseYear": 2023,
      "launchDate": "2023-09-06",
      "targetId": "gfx1101",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 624,
      "cacheMB": 64,
      "fp32TFLOPS": 37.3,
      "fp16TFLOPSVector": 37.3,
      "fp16TFLOPS": 74.6,
      "int8TOPS": 74.6,
      "int4TOPS": 149,
      "transistorCountBillion": 28.1,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 60,
      "shaderCoreCount": 3840,
      "matrixCoreCount": 120,
      "peakClockMHz": 2430,
      "tdpWatts": 263,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "msrpUSD": 499,
      "notes": "fp16TFLOPS (74.6, dense) is AMD's directly published \"Peak Half Precision (FP16 Matrix) Performance\" figure — no structured-sparsity variant listed for this card (RDNA3 doesn't get the 2:4-sparsity feature that RDNA4 parts do). fp8TFLOPS omitted — RDNA 3 has no native FP8 matrix support; AMD's page instead lists INT8/INT4 matrix rates (captured as int8TOPS/int4TOPS) via \"AI Accelerators\" (120, used as matrixCoreCount) rather than dedicated FP8 hardware. memoryBandwidthGBs uses the raw GDDR6 figure (624 GB/s); AMD's page also lists an \"Effective Memory Bandwidth\" of 2,708 GB/s (Infinity Cache-inflated), not used here to stay consistent with how every other entry in this collection reports raw, not cache-effective, bandwidth. No Launch Date field is published on this page; launchDate (2023-09-06) was added 2026-09-20 from AMD's own Gamescom press release of 2023-08-25, which states availability beginning September 6, 2023 for both this card and the RX 7700 XT (amd.com/en/newsroom/press-releases/2023-8-25-new-amd-radeon-rx-7800-xt-and-radeon-rx-7700-xt-gr.html). Source: amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7800-xt.html (\"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-radeon-rx-7900-gre",
      "vendor": "amd",
      "name": "Radeon RX 7900 GRE",
      "category": "consumer",
      "architecture": "RDNA 3",
      "releaseYear": 2023,
      "launchDate": "2023-07-27",
      "targetId": "gfx1100",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 576,
      "cacheMB": 64,
      "fp32TFLOPS": 46,
      "fp16TFLOPSVector": 46,
      "fp16TFLOPS": 92,
      "int8TOPS": 92,
      "int4TOPS": 184,
      "transistorCountBillion": 54,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 80,
      "shaderCoreCount": 5120,
      "matrixCoreCount": 160,
      "peakClockMHz": 2245,
      "tdpWatts": 260,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "A cut-down Navi 31 part, so it reports gfx1100 and inherits the RX 7900 XTX's ROCm support from 6.0.0 onward — useful to know, since the \"GRE\" naming gives no hint that it shares a target with the flagship.\ntransistorCountBillion is 54 B here versus 58 B for the 7900 XTX/XT on the same die family: AMD publishes the lower figure for this SKU, consistent with the GRE shipping fewer memory-cache dies (its 64 MB Infinity Cache and 256-bit interface versus the XTX's 96 MB / 384-bit). Recorded as published rather than normalised to the XTX's number.\nfp16TFLOPS (92, dense) is AMD's published \"Peak Half Precision (FP16 Matrix) Performance\"; the 46 TFLOPs vector rate is captured separately as fp16TFLOPSVector. No structured-sparsity variants are listed (RDNA 3 has no 2:4 sparsity). fp8TFLOPS omitted — no native FP8 matrix support on RDNA 3; INT8/INT4 matrix rates are published instead via the card's 160 \"AI Accelerators\" (used as matrixCoreCount).\nmemoryBandwidthGBs uses the raw GDDR6 figure (576 GB/s), not an Infinity Cache-effective number. peakClockMHz uses the published Boost Frequency (\"Up to 2245 MHz\"), not the 1880 MHz Game Frequency. AMD publishes no MSRP on this page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900-gre.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-7900-xt",
      "vendor": "amd",
      "name": "Radeon RX 7900 XT",
      "category": "consumer",
      "architecture": "RDNA 3",
      "releaseYear": 2022,
      "launchDate": "2022-12-13",
      "targetId": "gfx1100",
      "vramGB": 20,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 800,
      "cacheMB": 80,
      "fp32TFLOPS": 51.6,
      "fp16TFLOPSVector": 51.6,
      "fp16TFLOPS": 103,
      "int8TOPS": 103,
      "int4TOPS": 206,
      "transistorCountBillion": 58,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 84,
      "shaderCoreCount": 5376,
      "matrixCoreCount": 168,
      "peakClockMHz": 2400,
      "tdpWatts": 315,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Same Navi 31 die (gfx1100) as the RX 7900 XTX, so it shares that card's ROCm support from 6.0.0 onward. 20 GB of VRAM makes it the second-roomiest consumer Radeon in this collection after the XTX's 24 GB — relevant for local LLM work, where VRAM capacity binds before compute does.\nfp16TFLOPS (103, dense) is AMD's directly published \"Peak Half Precision (FP16 Matrix) Performance\" figure, distinct from the 51.6 TFLOPs vector rate captured as fp16TFLOPSVector. No structured-sparsity variant is listed (RDNA 3 lacks the 2:4-sparsity feature RDNA 4 parts have), so there is only one figure per precision to report. fp8TFLOPS omitted — RDNA 3 has no native FP8 matrix support; AMD's page instead publishes INT8/INT4 matrix rates (int8TOPS/int4TOPS) via its 168 \"AI Accelerators\" (used as matrixCoreCount).\nmemoryBandwidthGBs uses the raw GDDR6 figure (800 GB/s), not any Infinity Cache-inflated \"Effective Memory Bandwidth\" number, for consistency with the rest of this collection. peakClockMHz uses the published Boost Frequency (\"Up to 2400 MHz\"), not the 2000 MHz Game Frequency. AMD publishes no MSRP on this page, so msrpUSD is omitted. Note the URL has no hyphen before \"xt\" (same quirk as the 7900 XTX page). Source: amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xt.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-7900-xtx",
      "vendor": "amd",
      "name": "Radeon RX 7900 XTX",
      "category": "consumer",
      "architecture": "RDNA 3",
      "releaseYear": 2022,
      "launchDate": "2022-12-13",
      "targetId": "gfx1100",
      "vramGB": 24,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 960,
      "cacheMB": 96,
      "fp32TFLOPS": 61.4,
      "fp16TFLOPSVector": 61.4,
      "fp16TFLOPS": 123,
      "int8TOPS": 123,
      "int4TOPS": 246,
      "transistorCountBillion": 58,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 96,
      "shaderCoreCount": 6144,
      "matrixCoreCount": 192,
      "peakClockMHz": 2500,
      "tdpWatts": 355,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "msrpUSD": 999,
      "notes": "fp16TFLOPS (123) is directly published as \"Peak Half Precision (FP16 Matrix) Performance\" on AMD's own product page, not plain vector-shader FP32 (61.4 TFLOPs, much higher but not the matrix throughput number this collection uses; captured separately as fp16TFLOPSVector — also 61.4). No structured-sparsity variant is listed for this card (unlike RDNA4 parts), so there's only one figure to report. fp8TFLOPS omitted — RDNA 3 has no native FP8 matrix support (added in RDNA 4); AMD's page instead lists INT8/INT4 matrix rates (int8TOPS/int4TOPS) via \"AI Accelerators\" (192, used as matrixCoreCount). memoryBandwidthGBs uses the raw GDDR6 figure (960 GB/s), not the page's \"Effective Memory Bandwidth\" of 3,500 GB/s (Infinity Cache-inflated), for consistency with the rest of this collection. Source: amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xtx.html (\"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-radeon-rx-9060-xt",
      "vendor": "amd",
      "name": "Radeon RX 9060 XT (8GB)",
      "category": "consumer",
      "architecture": "RDNA 4",
      "releaseYear": 2025,
      "targetId": "gfx1200",
      "vramGB": 8,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 320,
      "cacheMB": 32,
      "fp32TFLOPS": 25.6,
      "fp16TFLOPSVector": 25.6,
      "fp16TFLOPS": 103,
      "fp8TFLOPS": 205,
      "int8TOPS": 205,
      "int4TOPS": 410,
      "fp16TFLOPSSparse": 205,
      "fp8TFLOPSSparse": 410,
      "int8TOPSSparse": 410,
      "transistorCountBillion": 29.7,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 32,
      "shaderCoreCount": 2048,
      "matrixCoreCount": 64,
      "peakClockMHz": 3130,
      "tdpWatts": 150,
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "fp16TFLOPS (103, dense) and fp8TFLOPS (205, dense) are both directly published on AMD's product page (\"Peak Half Precision (FP16 Matrix) Performance: 103 TFLOPs\", 205 with sparsity; \"Peak 8-bit Precision (FP8 Matrix) Performance: 205 TFLOPs\", 410 with sparsity) — RDNA4 is AMD's first Radeon architecture with native FP8 matrix support and 2:4 structured sparsity. int8TOPS/int8TOPSSparse mirror the FP8 figures exactly (205/410); int4TOPS listed but its structured-sparsity variant (821 TOPs) isn't captured in a separate schema field. matrixCoreCount uses the page's \"AI Accelerators\" count (64).\nAMD also sells a 16GB variant of this card under a separate SKU/URL that returned a 404 during an earlier data pull — only the 8GB variant (confirmed real, distinct VRAM/possibly binned clocks) is represented here; treat the 16GB variant as a coverage gap, not confirmed identical to this entry. No Launch Date field is published on this page, so launchDate is omitted; releaseYear (2025) is unaffected. Source: amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt-8gb.html (\"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-radeon-rx-9060-xt-16gb",
      "vendor": "amd",
      "name": "Radeon RX 9060 XT (16GB)",
      "category": "consumer",
      "architecture": "RDNA 4",
      "releaseYear": 2025,
      "targetId": "gfx1200",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 320,
      "cacheMB": 32,
      "fp32TFLOPS": 25.6,
      "fp16TFLOPSVector": 25.6,
      "fp16TFLOPS": 103,
      "fp8TFLOPS": 205,
      "int8TOPS": 205,
      "int4TOPS": 410,
      "fp16TFLOPSSparse": 205,
      "fp8TFLOPSSparse": 410,
      "int8TOPSSparse": 410,
      "transistorCountBillion": 29.7,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 32,
      "shaderCoreCount": 2048,
      "matrixCoreCount": 64,
      "peakClockMHz": 3130,
      "tdpWatts": 160,
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Closes the coverage gap flagged in the RX 9060 XT (8GB) entry, whose notes recorded that the 16 GB SKU's URL 404'd during the 2026-09-01 pull. AMD serves the 16 GB variant at the unsuffixed URL (amd-radeon-rx-9060xt.html, page title \"AMD Radeon RX 9060 XT (16GB)\") while the 8 GB variant lives at amd-radeon-rx-9060xt-8gb.html — worth remembering, since guessing a \"-16gb\" suffix returns a 404.\nCompute specs are identical to the 8 GB variant across the board (32 CUs, 2048 stream processors, 64 AI accelerators, 3130 MHz boost, same FP32/FP16/FP8/INT8/INT4 figures) — the two SKUs differ only in memory capacity (16 GB vs 8 GB) and board power (160 W vs 150 W). Memory bandwidth is the same 320 GB/s on a 128-bit bus at 20 Gbps, so the 16 GB card fits larger models without feeding them any faster. For local LLM use that still makes it the more sensible of the two: 8 GB is the binding constraint on the smaller SKU well before compute is.\nShares gfx1200 (Navi 44) with the 8 GB variant, so ROCm support is the same — from 6.4.1 (2025-05) onward.\nfp16TFLOPS (103, dense) and fp8TFLOPS (205, dense) are both directly published with explicit \"with Structured Sparsity\" variants (205 / 410) recorded in the ...Sparse fields. int8TOPS/int8TOPSSparse mirror the FP8 figures exactly (205/410); int4TOPS is listed but its sparsity variant (821 TOPs) has no schema field. matrixCoreCount uses the page's \"AI Accelerators\" count (64).\npeakClockMHz uses the published Boost Frequency (\"Up to 3130 MHz\"), not the 2530 MHz Game Frequency. No Launch Date field is published on this page and AMD's COMPUTEX 2025 announcement (2025-05-20) gives no specific availability date, so launchDate is omitted; releaseYear (2025) is unaffected. AMD publishes no MSRP on the spec page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt.html (\"Expand All\" spec accordion), pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-9070",
      "vendor": "amd",
      "name": "Radeon RX 9070",
      "category": "consumer",
      "architecture": "RDNA 4",
      "releaseYear": 2025,
      "launchDate": "2025-03-06",
      "targetId": "gfx1201",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 640,
      "cacheMB": 64,
      "fp32TFLOPS": 36.1,
      "fp16TFLOPSVector": 36.1,
      "fp16TFLOPS": 145,
      "fp8TFLOPS": 289,
      "int8TOPS": 289,
      "int4TOPS": 578,
      "fp16TFLOPSSparse": 289,
      "fp8TFLOPSSparse": 578,
      "int8TOPSSparse": 578,
      "transistorCountBillion": 53.9,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 56,
      "shaderCoreCount": 3584,
      "matrixCoreCount": 112,
      "peakClockMHz": 2520,
      "tdpWatts": 220,
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Same Navi 48 die (gfx1201) as the RX 9070 XT, so it shares that card's ROCm support from 6.4.1 (2025-05) onward — roughly two and a half months after launch, the fastest launch-to-support turnaround of any consumer Radeon in this collection.\nfp16TFLOPS (145, dense) and fp8TFLOPS (289, dense) are both directly published, each with an explicit \"with Structured Sparsity\" variant (289 / 578) recorded in the ...Sparse fields — AMD publishes the dense figure as the headline and sparsity separately, the opposite of NVIDIA's convention, so no halving is applied here. int8TOPS/int8TOPSSparse mirror the FP8 figures exactly (289/578); int4TOPS is listed but its sparsity variant (1,156 TOPs) has no schema field. matrixCoreCount uses the page's \"AI Accelerators\" count (112).\nSame 16 GB / 640 GB/s memory configuration as the 9070 XT, at lower clocks and 220 W instead of 304 W — for VRAM-bound local LLM work the two cards are closer than their compute figures suggest.\npeakClockMHz uses the published Boost Frequency (\"Up to 2520 MHz\"), not the 2070 MHz Game Frequency. tdpWatts is the published \"Typical Board Power (Desktop)\". AMD's product page publishes no Launch Date field; launchDate (2025-03-06) comes from AMD's own RDNA 4 launch press release of 2025-02-28, which states availability beginning March 6, 2025. AMD publishes no MSRP on the spec page, so msrpUSD is omitted. Source: amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070.html (\"Expand All\" spec accordion) plus amd.com/en/newsroom/press-releases/2025-2-28-amd-unveils-next-generation-amd-rdna-4-architectu.html for the date, pulled 2026-09-20.\n"
    },
    {
      "id": "amd-radeon-rx-9070-xt",
      "vendor": "amd",
      "name": "Radeon RX 9070 XT",
      "category": "consumer",
      "architecture": "RDNA 4",
      "releaseYear": 2025,
      "launchDate": "2025-03-06",
      "targetId": "gfx1201",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 640,
      "cacheMB": 64,
      "fp32TFLOPS": 48.7,
      "fp16TFLOPSVector": 48.7,
      "fp16TFLOPS": 195,
      "fp8TFLOPS": 389,
      "int8TOPS": 389,
      "int4TOPS": 779,
      "fp16TFLOPSSparse": 389,
      "fp8TFLOPSSparse": 779,
      "int8TOPSSparse": 779,
      "transistorCountBillion": 53.9,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 64,
      "shaderCoreCount": 4096,
      "matrixCoreCount": 128,
      "peakClockMHz": 2970,
      "tdpWatts": 304,
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "msrpUSD": 599,
      "notes": "Re-verified directly against AMD's own product page (the correct URL is amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html, no hyphen before \"xt\"). fp16TFLOPS (195, dense) and fp8TFLOPS (389, dense) are both directly published (\"Peak Half Precision (FP16 Matrix) Performance: 195 TFLOPs\", 389 with structured sparsity; \"Peak 8-bit Precision (FP8 Matrix) Performance: 389 TFLOPs\", 779 with sparsity) — RDNA4 is AMD's first Radeon architecture with native FP8 matrix support. int8TOPS/int8TOPSSparse mirror the FP8 figures exactly (389/779); int4TOPS listed but its sparsity variant (1,557 TOPs) isn't captured in a separate schema field. matrixCoreCount uses the page's \"AI Accelerators\" count (128). tdpWatts (304W, \"Typical Board Power\") and memoryBandwidthGBs (640) match this file's prior corroborated figures exactly. No Launch Date field is published on this page; launchDate (2025-03-06) was added 2026-09-20 from AMD's own RDNA 4 launch press release of 2025-02-28, which states availability beginning March 6, 2025 for both this card and the RX 9070 (amd.com/en/newsroom/press-releases/2025-2-28-amd-unveils-next-generation-amd-rdna-4-architectu.html). Source: amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html (\"Expand All\" spec accordion), pulled 2026-09-01.\n"
    },
    {
      "id": "amd-radeon-vii",
      "vendor": "amd",
      "name": "Radeon VII",
      "category": "consumer",
      "architecture": "GCN 5.1 (Vega 20)",
      "releaseYear": 2019,
      "targetId": "gfx906",
      "vramGB": 16,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 1024,
      "fp32TFLOPS": 13.8,
      "fp16TFLOPSVector": 27.7,
      "transistorCountBillion": 13.2,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 60,
      "shaderCoreCount": 3840,
      "peakClockMHz": 1750,
      "tdpWatts": 300,
      "interconnect": "PCIe only (no Infinity Fabric multi-GPU link)",
      "notes": "Added 2026-09-21 after an r/ROCm commenter asked \"no gfx906?\" — the collection had the Instinct MI50 (also gfx906) but not the consumer Vega 20 card, which is the one people actually encounter second-hand.\nThe most important fact about this card is that its support was REMOVED. gfx906 was supported from ROCm 4.5.0 through 6.3.3, listed as Deprecated (still enabled, removal announced) from 6.0.0 on, and dropped in 6.4.0 (2025-04-11). Until 2026-09-27 this note said 6.0.2 / 6.1.0: AMD's compatibility-matrix page omits deprecated targets, and its install-on-linux system-requirements page is where 6.1–6.3 list them. A Reddit commenter caught it. AMD's current 6.4.0 page shows gfx906 as Unsupported, but the Wayback Machine copy from 2025-05-13 lists the Radeon VII and Radeon PRO VII as Deprecated there too, so the page was changed after release. It is the only consumer ROCm gfx target in this collection that ever lost support — every other gfx target either gained it and kept it, or never had it. That is a statement about ROCm targets specifically, NOT about this collection as a whole: NVIDIA has retired Kepler, Maxwell, Pascal and Volta, which current CUDA (compute capability 7.5 and up) no longer lists, covering more cards here than the AMD side does. See /rocm-consumer-gpus/ for the derived comparison. So this card reads as \"supported\" in any guide, blog post or Stack Overflow answer written before April 2025, and is a dead end on every currently-supported ROCm release. See /rocm-consumer-gpus/ for the options.\nWhy that matters more than it sounds: 16 GB of HBM2 at 1024 GB/s is still more memory bandwidth than any consumer Radeon released since, including the RX 7900 XTX (960 GB/s), on a card that now sells cheaply second-hand. On paper it looks like the bargain of the local-LLM world. The bandwidth is real; the software support is gone.\nfp16TFLOPS and fp8TFLOPS intentionally omitted — Vega has no matrix/AI-accelerator cores (AMD's spec record publishes no \"AI Accelerators\" count for this card, unlike RDNA 3/4 parts), so the 27.7 TFLOPs figure is the packed-math FP16 vector rate, captured as fp16TFLOPSVector. matrixCoreCount omitted for the same reason.\nfp64TFLOPS is deliberately NOT recorded. This card is well known for an unusually generous 1:4 FP64 rate for a consumer part, but AMD's own spec record publishes no double-precision figure for it, so nothing is entered here rather than deriving it from the FP32 number. cacheMB omitted — Vega predates AMD Infinity Cache entirely, so there is no comparable figure (not a data gap). peakClockMHz uses the published Boost Frequency (1750 MHz); base is 1400 MHz.\nAMD's product page for this card is gone (every URL pattern tried returns 404), so the data comes from AMD's own live specifications comparison tool at amd.com/en/products/specifications/graphics.html, which still carries a structured spec record for it in an embedded JSON data model — the same sourcing route used for the MI6/MI8 entries. That record publishes Launch Date as the bare year \"2019\" with no month or day, so launchDate is omitted and only releaseYear is set. No MSRP, no PCIe bus type and no ECC line are published there either. Source: amd.com/en/products/specifications/graphics.html (embedded product record \"AMDRadeonVII\"), pulled 2026-09-21.\n"
    },
    {
      "id": "amd-ryzen-ai-max-plus-395",
      "vendor": "amd",
      "name": "Ryzen AI Max+ 395",
      "category": "consumer",
      "architecture": "Radeon 8060S (Strix Halo)",
      "releaseYear": 2025,
      "targetId": "gfx1151",
      "vramGB": 128,
      "memoryType": "LPDDR5x-8000 (unified, shared with the CPU)",
      "memoryBandwidthGBs": 256,
      "computeUnitLabel": "Compute Units",
      "computeUnitCount": 40,
      "peakClockMHz": 2900,
      "tdpWatts": 55,
      "notes": "Added 2026-09-27: the \"Strix Halo\" APU whose integrated Radeon 8060S can use up to 128 GB of system memory, popular for local LLMs in mini PCs and laptops. Sources: AMD's product page (amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html): Radeon 8060S Graphics, 40 graphics cores, 2,900 MHz, 256-bit LPDDR5x, LPDDR5x-8000, max 128 GB, default TDP 55 W (cTDP 45-120 W, so the same chip runs at very different power in different machines), former codename Strix Halo. memoryBandwidthGBs is derived: 8,000 MT/s × 256 bits ÷ 8 = 256 GB/s. vramGB is the maximum system memory; how much the GPU may use depends on the machine's firmware and OS settings, and the CPU shares it. gfx1151 and ROCm support from AMD's ROCm 7.14.0 / 10.0.0 compatibility matrices (Ryzen APU section, \"AMD Ryzen AI Max+ 395 (Radeon 8060S) (gfx1151)\"). releaseYear 2025: AMD announced the Ryzen AI Max series at CES in January 2025; the product page gives no launch date. AMD doesn't name the GPU architecture on this page, so none is claimed beyond the graphics model.\n"
    },
    {
      "id": "intel-arc-a380",
      "vendor": "intel",
      "name": "Arc A380",
      "category": "consumer",
      "architecture": "Xe-HPG (Alchemist)",
      "releaseYear": 2022,
      "launchDate": "Aug 2022 (Q2'22 per ark.intel.com; first US retail availability ~Aug 22, 2022)",
      "vramGB": 6,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 186,
      "int8TOPS": 66,
      "processNode": "TSMC N6",
      "computeUnitLabel": "Xe-cores",
      "computeUnitCount": 8,
      "shaderCoreCount": 128,
      "matrixCoreCount": 128,
      "peakClockMHz": 2000,
      "tdpWatts": 75,
      "pcieGen": "PCIe 4.0 x8",
      "msrpUSD": 139,
      "notes": "Added 2026-09 alongside Arc A750 — Intel's smallest/cheapest Alchemist card, notable in the field mainly for AV1 hardware encode at a very low price point rather than ML throughput. Source: ark.intel.com product specifications page for the Arc A380 (intel.com/.../products/sku/227959/.../specifications.html), pulled 2026-09 — same sourcing tier as this collection's A770/A750 entries. shaderCoreCount (128) is ark's \"Xe Vector Engines\" count and matrixCoreCount (128) is its \"Intel XMX Engines\" count — structural counts, not throughput figures, matching those entries' convention. pcieGen is \"PCIe 4.0 x8\" specifically (not x16) — ark states \"Up to PCI Express 4.0 x8 (x16 slot required)\", a real bandwidth-halving distinction from the larger Arc cards in this collection, not a typo.\nfp16TFLOPS/bf16TFLOPS/fp8TFLOPS are intentionally omitted, same reasoning as A770/A750: Intel's ark page for Arc consumer cards doesn't publish a per-precision matrix/tensor FLOPS breakdown, only the single \"GPU Peak TOPS (Int8)\" figure used here (66). msrpUSD (139) is the original Aug 2022 US retail launch price, well-corroborated across contemporaneous outlets (Neowin, TechPowerUp, WCCFTech), not itself listed on the current ark page.\n"
    },
    {
      "id": "intel-arc-a750",
      "vendor": "intel",
      "name": "Arc A750",
      "category": "consumer",
      "architecture": "Xe-HPG (Alchemist)",
      "releaseYear": 2022,
      "launchDate": "Oct 12, 2022 (Q3'22 per ark.intel.com)",
      "vramGB": 8,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 512,
      "int8TOPS": 229,
      "processNode": "TSMC N6",
      "computeUnitLabel": "Xe-cores",
      "computeUnitCount": 28,
      "shaderCoreCount": 448,
      "matrixCoreCount": 448,
      "peakClockMHz": 2050,
      "tdpWatts": 225,
      "pcieGen": "PCIe 4.0 x16",
      "msrpUSD": 289,
      "notes": "Added 2026-09 alongside Arc A380 to round out this collection's Alchemist coverage (previously only the flagship A770 was represented). Source: ark.intel.com product specifications page for the Arc A750 (intel.com/.../products/sku/227954/.../specifications.html), pulled 2026-09 — same page style and sourcing tier as this collection's existing A770 entry. shaderCoreCount (448) is ark's \"Xe Vector Engines\" count and matrixCoreCount (448) is its \"Intel XMX Engines\" count — both structural counts, not throughput figures, same convention as A770.\nfp16TFLOPS/bf16TFLOPS/fp8TFLOPS are intentionally omitted, same reasoning as A770: Intel's ark page for Arc consumer cards doesn't publish a per-precision matrix/tensor FLOPS breakdown, only the single \"GPU Peak TOPS (Int8)\" figure used here (229). msrpUSD (289) is the original Oct 12 2022 launch price — well-corroborated across contemporaneous outlets (PC Gamer, TechSpot, WCCFTech all cite the same Intel-announced $289 figure), not itself listed on the current ark page (which shows no pricing).\n"
    },
    {
      "id": "intel-arc-a770",
      "vendor": "intel",
      "name": "Arc A770 (16GB)",
      "category": "consumer",
      "architecture": "Xe-HPG (Alchemist)",
      "releaseYear": 2022,
      "launchDate": "Q3'22",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 560,
      "int8TOPS": 262,
      "processNode": "TSMC N6",
      "computeUnitLabel": "Xe-cores",
      "computeUnitCount": 32,
      "shaderCoreCount": 512,
      "matrixCoreCount": 512,
      "peakClockMHz": 2100,
      "tdpWatts": 225,
      "pcieGen": "PCIe 4.0 x16",
      "msrpUSD": 349,
      "notes": "Source: ark.intel.com product specifications page for the Arc A770 16GB SKU (intel.com/.../products/sku/229151/.../specifications.html), pulled 2026-09-01 (a prior note in this file claiming ark.intel.com was blocked with HTTP 403 was from an earlier research pass using a plain fetch tool without JS/browser rendering — a direct browser visit works fine).\nfp16TFLOPS/bf16TFLOPS/fp8TFLOPS are still intentionally omitted — Intel's ark page for Arc consumer cards does not publish a per-precision matrix/tensor FLOPS breakdown at all, only the single \"GPU Peak TOPS (Int8)\" figure used here. shaderCoreCount (512) is ark's \"Xe Vector Engines\" count and matrixCoreCount (512) is its \"Intel XMX Engines\" count — both structural counts, not throughput figures.\nmsrpUSD (349) is Intel's original 2022 launch-day price; ark.intel.com no longer lists a \"Recommended Customer Price\" for this SKU (unlike newer Arc parts such as the B580), so this is carried over from the original announcement rather than the current ark page.\n"
    },
    {
      "id": "intel-arc-b580",
      "vendor": "intel",
      "name": "Arc B580",
      "category": "consumer",
      "architecture": "Xe2-HPG (Battlemage)",
      "releaseYear": 2024,
      "launchDate": "Q4'24",
      "vramGB": 12,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 456,
      "int8TOPS": 233,
      "processNode": "TSMC N5",
      "computeUnitLabel": "Xe-cores",
      "computeUnitCount": 20,
      "shaderCoreCount": 160,
      "matrixCoreCount": 160,
      "peakClockMHz": 2670,
      "tdpWatts": 190,
      "pcieGen": "PCIe 4.0 x8",
      "msrpUSD": 249,
      "notes": "Source: ark.intel.com product specifications page for the Arc B580 (intel.com/.../products/sku/241598/.../specifications.html), pulled 2026-09-01 (a prior note in this file claiming ark.intel.com was blocked with HTTP 403 was from an earlier research pass using a plain fetch tool without JS/browser rendering — a direct browser visit works fine).\nfp16TFLOPS/bf16TFLOPS/fp8TFLOPS are still intentionally omitted — Intel's ark page for Arc consumer cards does not publish a per-precision matrix/tensor FLOPS breakdown at all, only the single \"GPU Peak TOPS (Int8)\" figure used here. shaderCoreCount (160) is ark's \"Xe Vector Engines\" count and matrixCoreCount (160) is its \"Intel XMX Engines\" count — both structural counts, not throughput figures.\n"
    },
    {
      "id": "intel-data-center-gpu-max-1550",
      "vendor": "intel",
      "name": "Data Center GPU Max 1550",
      "category": "datacenter",
      "architecture": "Xe-HPC (Ponte Vecchio)",
      "releaseYear": 2023,
      "launchDate": "Q1'23",
      "vramGB": 128,
      "memoryType": "HBM2e",
      "memoryBandwidthGBs": 3276.8,
      "computeUnitLabel": "Xe-cores",
      "computeUnitCount": 128,
      "shaderCoreCount": 1024,
      "matrixCoreCount": 1024,
      "peakClockMHz": 1600,
      "tdpWatts": 600,
      "pcieGen": "PCIe 5.0 x16",
      "interconnect": "Xe Link",
      "notes": "Source: ark.intel.com product specifications page (intel.com/.../products/sku/232873/.../specifications.html), pulled 2026-09-01 (a prior note in this file claiming ark.intel.com was blocked with HTTP 403 was from an earlier research pass using a plain fetch tool without JS/browser rendering — a direct browser visit works fine).\nfp64TFLOPS/fp32TFLOPS/fp16TFLOPS/bf16TFLOPS are intentionally omitted — ark.intel.com does not publish a per-precision FLOPS breakdown for this part at all (unlike AMD/NVIDIA datacenter parts), and Intel's own Data Center GPU Max Series product brief PDF (the one document that likely does have these numbers) returned HTTP 403 via fetch and rendered as a blank/undecodable canvas via direct browser navigation — couldn't be read either way this session. The previous version of this file carried a \"derived\" fp16TFLOPS (2x an unverified 52 FP32 TFLOPS figure) — removed, since neither number could be confirmed from an official Intel source and this site doesn't publish self-derived estimates. Widely-cited secondary figures (~52 FP32/FP64 TFLOPS-class, ~13-team sources disagreeing by 30%+) exist but are not used here per this site's primary-source-only policy.\nmatrixCoreCount/shaderCoreCount (1024 each) are ark's \"Intel XMX Engines\" and \"Xe Vector Engines\" counts — structural counts, not throughput figures; ark doesn't publish a \"GPU Peak TOPS\" figure for this SKU either (unlike Arc consumer parts), so int8TOPS is also omitted.\ninterconnect is \"Xe Link\" (multi-GPU fabric for OAM baseboards); ark publishes only \"Intel Xe Link Maximum Frequency: 53 Gbps\" (a per-link signaling rate), not an aggregate bandwidth figure comparable to AMD's Infinity Fabric or NVIDIA's NVLink GB/s numbers, so interconnectBandwidthGBs is left blank rather than converting/guessing.\nNo targetId: Intel doesn't publish a clean oneAPI/Level Zero device-architecture identifier comparable to CUDA compute capability or ROCm gfx targets, so there's no compatibility cross-link for Intel GPUs yet.\nark.intel.com lists \"Expected Discontinuance: Jan 2026\" for this SKU — as of this update (Sep 2026) this part may already be past its end-of-life date; Intel's successor roadmap for this line (Rialto Bridge/Falcon Shores) remains unclear.\n"
    },
    {
      "id": "nvidia-a100-sxm4-80gb",
      "vendor": "nvidia",
      "name": "A100 SXM4 80GB",
      "category": "datacenter",
      "architecture": "Ampere",
      "releaseYear": 2020,
      "targetId": "8.0",
      "vramGB": 80,
      "memoryType": "HBM2e",
      "memoryBandwidthGBs": 2039,
      "fp64TFLOPS": 9.7,
      "fp32TFLOPS": 19.5,
      "fp64TFLOPSMatrix": 19.5,
      "tf32TFLOPS": 156,
      "bf16TFLOPS": 312,
      "fp16TFLOPS": 312,
      "int8TOPS": 624,
      "fp16TFLOPSSparse": 624,
      "int8TOPSSparse": 1248,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 108,
      "shaderCoreCount": 6912,
      "matrixCoreCount": 432,
      "tdpWatts": 400,
      "formFactor": "SXM",
      "pcieGen": "PCIe Gen4",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 600,
      "multiInstanceSupport": "Up to 7 MIGs @ 10GB each",
      "notes": "fp8TFLOPS intentionally omitted — Ampere's Tensor Cores don't support native FP8; that was introduced with Hopper. All figures sourced directly from NVIDIA's official A100 datasheet (nvidia-a100-datasheet-nvidia-us-2188504-web.pdf, \"80GB SXM\" column), pulled 2026-09-01 via pdftotext extraction of the PDF (WebFetch could not parse this PDF's text layer directly).\nfp64TFLOPS/fp64TFLOPSMatrix, fp32TFLOPS, tf32TFLOPS, bf16TFLOPS, fp16TFLOPS, and int8TOPS are all the vendor's own DENSE (non-sparse) figures — the datasheet lists each Tensor Core row as \"X TFLOPS | Y TFLOPS*\" with footnote \"* With sparsity\", so X is dense and Y is the sparsity figure. Only fp16 and int8 have a dedicated Sparse field in this schema (fp16TFLOPSSparse: 624, int8TOPSSparse: 1248); TF32's sparse figure is 312 TFLOPS and BF16's is 624 TFLOPS but there's no tf32/bf16 Sparse field — noted here instead.\ntdpWatts (400) is the standard SXM4 config; the datasheet notes an HGX A100-80GB CTS (Custom Thermal Solution) SKU can go up to 500W. interconnectBandwidthGBs (600) is NVLink; the SXM module also exposes a PCIe Gen4 x16 host link at 64GB/s (pcieGen field records the generation, not this secondary bandwidth figure). multiInstanceSupport confirms MIG partitioning into up to 7 instances @ 10GB VRAM each. No msrpUSD — A100 ships through OEM/server partners, not direct retail. No transistor count, L2 cache size, or process node found in the datasheet itself (only in NVIDIA's longer Ampere architecture whitepaper, not cross-checked here) — omitted rather than sourced from a secondary site.\n2026-09-15 — added computeUnitCount (108 SMs), shaderCoreCount (6,912 FP32 CUDA cores) and matrixCoreCount (432 third-generation Tensor Cores) from NVIDIA's own \"NVIDIA Ampere Architecture In-Depth\" developer blog post, which states the A100 product configuration of GA100 directly: \"108 SMs; 64 FP32 CUDA Cores/SM, 6912 FP32 CUDA Cores per GPU; 4 third-generation Tensor Cores/SM, 432 third-generation Tensor Cores per GPU.\" These counts are identical across every A100 SKU (40GB/80GB, SXM4/PCIe) — only memory, bandwidth and power differ between them.\n"
    },
    {
      "id": "nvidia-b200",
      "vendor": "nvidia",
      "name": "B200",
      "category": "datacenter",
      "architecture": "Blackwell",
      "releaseYear": 2024,
      "targetId": "10.0",
      "vramGB": 180,
      "memoryType": "HBM3e",
      "memoryBandwidthGBs": 7700,
      "fp64TFLOPS": 37,
      "fp32TFLOPS": 75,
      "fp64TFLOPSMatrix": 37,
      "tf32TFLOPS": 1100,
      "bf16TFLOPS": 2250,
      "fp16TFLOPS": 2250,
      "fp8TFLOPS": 4500,
      "fp6TFLOPS": 4500,
      "fp4TFLOPS": 9000,
      "int8TOPS": 4500,
      "fp16TFLOPSSparse": 4500,
      "fp8TFLOPSSparse": 9000,
      "int8TOPSSparse": 9000,
      "processNode": "TSMC 4NP",
      "transistorCountBillion": 208,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 148,
      "tdpWatts": 1000,
      "formFactor": "SXM",
      "pcieGen": "PCIe Gen5",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 1800,
      "multiInstanceSupport": "Up to 7 MIG instances",
      "notes": "Superseded 2026-09-01: this file previously derived per-GPU figures by dividing 8-GPU HGX B200 baseboard totals by 8. NVIDIA's official \"NVIDIA Blackwell\" architecture datasheet (linked from nvidia.com/en-us/data-center/hgx/ as \"Read the NVIDIA Blackwell Datasheet\", PDF, pulled 2026-09-01 via pdftotext) turns out to publish an \"Individual Blackwell GPU Specifications\" table directly — i.e. true per-GPU numbers, not something this site had to derive. All compute/ memory/TDP/interconnect fields below are read directly from that table's HGX B200 column, not derived. (NVIDIA's B200-specific product page at nvidia.com/en-us/data-center/b200/ now redirects to the general HGX platform page as of this pull — Blackwell has been superseded by Rubin in NVIDIA's current lineup.)\ntf32TFLOPS, bf16TFLOPS, fp16TFLOPS, fp8TFLOPS, fp6TFLOPS, fp4TFLOPS, and int8TOPS are dense figures — the datasheet states each row as sparse, with \"dense is one-half of the sparse spec shown\" (its own footnote 2). The published sparse figures per GPU: FP4 18 PFLOPS, FP8/FP6 9 PFLOPS (one shared row — FP6 and FP8 have identical throughput on this part), INT8 9 POPS, FP16/BF16 4.5 PFLOPS, TF32 2.2 PFLOPS — halved here for the dense fields; fp16TFLOPSSparse/fp8TFLOPSSparse/int8TOPSSparse hold the sparse figures directly (no dedicated sparse field exists in this schema for tf32/fp4/fp6, so their sparse values — 2200/18000/18000 TFLOPS respectively — are recorded here in notes only). FP32 (75 TFLOPS) and FP64/FP64 Tensor Core (37 TFLOPS, one combined row — no separate vector/matrix split published) are NOT marked with the sparsity footnote and are used as-is.\nmemoryBandwidthGBs corrected from 7750 to 7700 to match the datasheet's precise \"180 GB HBM3E | 8 TB/s\" per-GPU... actually the exact line reads \"180 GB HBM3E | 7.7 TB/s\" for the HGX B200 column — 7700 GB/s used accordingly (the previous 7750 figure came from the old /8-derivation method and is superseded).\ntransistorCountBillion (208) and processNode (\"TSMC 4NP\") describe the full dual-die Blackwell package as one GPU, per the datasheet's own architecture description (\"NVIDIA Blackwell architecture GPUs pack 208 billion transistors and are manufactured using a custom-built TSMC 4NP process... two reticle-limited dies connected by a 10 TB/s chip-to-chip interconnect in a unified single GPU\").\nmultiInstanceSupport: the datasheet's MIG row simply states \"7\" (spanning all three server configs in that table) rather than a per-partition VRAM size like Hopper's \"@10GB each\" — recorded as-is rather than computing an unstated partition size. No msrpUSD — ships through OEM/server partners.\n2026-09-15 — added computeUnitCount (148 SMs), flagged here as the one figure in this file NOT traceable to an NVIDIA-published document. NVIDIA's B200 product materials omit the SM count entirely. 148 — 74 of the 80 physically present SMs enabled per die, across two dies — is reported by chipsandcheese.com's B200 architecture deep-dive, which attributes it to NVIDIA's own Hot Chips 2024 Blackwell presentation and corroborates it against hands-on testing of real B200 hardware. It is consistent with NVIDIA's Blackwell Ultra blog putting the full, nothing-fused-off two-die package at 160 SMs (see this collection's B300 entry): B200 is the partially-harvested bin of the same silicon.\nCaveat worth keeping for anyone re-checking this: some third-party references (e.g. Cornell CAC's GPU-architecture pages) state 160 SMs for B200, but they cite the Blackwell *Ultra* blog as the source for it — that looks like the Ultra figure applied to the wrong SKU rather than an independent measurement, so it isn't treated here as a competing source. shaderCoreCount and matrixCoreCount are deliberately left empty rather than derived as 148 x 128 and 148 x 4: the per-SM structure is right, but multiplying an already-secondhand SM count would dress a doubly-derived number up as vendor data.\n"
    },
    {
      "id": "nvidia-b300",
      "vendor": "nvidia",
      "name": "B300 (Blackwell Ultra)",
      "category": "datacenter",
      "architecture": "Blackwell Ultra",
      "releaseYear": 2026,
      "targetId": "10.3",
      "vramGB": 288,
      "memoryType": "HBM3e",
      "memoryBandwidthGBs": 8000,
      "fp64TFLOPS": 1.25,
      "fp32TFLOPS": 75,
      "fp64TFLOPSMatrix": 1.25,
      "tf32TFLOPS": 1125,
      "bf16TFLOPS": 2250,
      "fp16TFLOPS": 2250,
      "fp8TFLOPS": 4500,
      "fp6TFLOPS": 4500,
      "fp4TFLOPS": 13500,
      "fp16TFLOPSSparse": 4500,
      "fp8TFLOPSSparse": 9000,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 160,
      "shaderCoreCount": 20480,
      "matrixCoreCount": 640,
      "tdpWatts": 1400,
      "formFactor": "SXM",
      "pcieGen": "PCIe Gen5",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 1800,
      "notes": "Newly launched (NVIDIA's own DGX B300 page: \"NVIDIA DGX B300 Systems Are Shipping Now\") — added alongside AMD's MI350X/MI355X/MI455X as the current-generation refresh this collection was missing. B300's compute capability is sm_103 (distinct from B200's sm_100). Until 2026-09-27 the targetId was \"10.0\", because the cuda collection had no 10.3 entry to link to; CUDA 12.9's release notes add sm_103, so the cuda data now lists 10.3 from 12.9 on and targetId is the real value. Library and PyTorch verdicts don't change: sm_100 code runs on sm_103 under CUDA's same-major rule.\nUpdated 2026-09-01: per-GPU compute figures now sourced from NVIDIA's HGX B300/HGX B200 comparison table (nvidia.com/en-us/data-center/hgx/, \"HGX Specifications\" section, HGX B300 column), which gives true 8-GPU totals per precision that this site derives to per-GPU by /8 — the standalone Blackwell Ultra PDF datasheet (linked from that page) is behind a lead-gen form (resources.nvidia.com) and wasn't accessible. fp4TFLOPS/fp6TFLOPS are new fields not previously captured; fp8TFLOPS/ fp16TFLOPS are unchanged from the prior derivation (confirmed independently against this new source — identical 4,500/2,250 dense values).\nAll values below are dense (non-sparse); NVIDIA's table states each row as sparse with footnote \"Dense is ½ sparse spec shown\" except FP32 and FP64. 8-GPU sparse|dense totals: FP4 144|108 PFLOPS, FP8/FP6 (one shared row — same throughput for both) 72|36 PFLOPS, FP16/BF16 36|18 PFLOPS, TF32 18|9 PFLOPS — divided by 8 for the per-GPU dense figures here (fp16TFLOPSSparse/fp8TFLOPSSparse hold the per-GPU sparse values; no sparse field exists in this schema for tf32/fp4/fp6 — their per-GPU sparse values are 2,250/18,000/9,000 TFLOPS respectively, noted here only). FP32 (600 TFLOPS/8 = 75) and FP64/FP64 Tensor Core (10 TFLOPS/8 = 1.25) are not sparsity-marked and used as-is.\nint8TOPS intentionally omitted: NVIDIA's table lists the 8-GPU INT8 figure as \"3 POPS\" (sparse) — roughly 1/24th of the FP8 row's 72 PFLOPS, wildly inconsistent with every other generation in this collection where INT8 and FP8 throughput match closely. This looks like an error or unit typo in NVIDIA's own published table (possibly missing a digit) rather than a real architectural cut, so it's left out rather than propagated as fact — worth rechecking against a future NVIDIA source.\nfp64TFLOPS (1.25 dense per GPU) is a genuine, dramatic cut from B200's 37 TFLOPS — directionally consistent with public reporting that Blackwell Ultra reallocates die area from FP64/HPC toward FP4 inference throughput, and unlike the INT8 figure this one is internally plausible (not an outlier relative to the rest of the row), so it's kept.\nvramGB (288) and memoryBandwidthGBs (8000) remain corroborated across independent secondary sources rather than NVIDIA's own quick-specs page directly — that page states only an aggregate \"Total GPU Memory: 2.1 TB\" for the 8-GPU system (262.5GB/GPU if evenly divided), which doesn't match the widely-reported 288GB/GPU figure; likely a usable-vs-raw- capacity difference (same kind of discrepancy seen on this site's B200 entry). tdpWatts (1,400) remains a well-corroborated secondary-source figure, not a single NVIDIA-published per-GPU number. pcieGen (PCIe Gen5) is inferred from the general Blackwell-family interconnect spec (shared across B200/B300 per NVIDIA's architecture materials), not a B300-specific row in the comparison table. No transistor count or process node included — not confirmed specifically for the Ultra variant (only for baseline Blackwell/B200) and not assumed. No msrpUSD — ships through OEM/server partners.\n2026-09-15 — added computeUnitCount (160 SMs), shaderCoreCount (20,480 CUDA cores) and matrixCoreCount (640 fifth-generation Tensor Cores) from NVIDIA's own \"Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era\" developer blog post, which states \"160 Streaming Multiprocessors (SMs) organized into eight Graphics Processing Clusters\". 160 is the full two-die Blackwell package with nothing fused off (80 SMs per die) — which is precisely what separates Blackwell Ultra's die configuration from B200's partially-harvested 148, see this collection's B200 entry.\n"
    },
    {
      "id": "nvidia-dgx-spark",
      "vendor": "nvidia",
      "name": "DGX Spark",
      "category": "workstation",
      "architecture": "Blackwell (GB10)",
      "releaseYear": 2025,
      "targetId": "12.1",
      "vramGB": 128,
      "memoryType": "LPDDR5x (unified, shared with the CPU)",
      "memoryBandwidthGBs": 273,
      "fp4TFLOPS": 500,
      "tdpWatts": 140,
      "notes": "Added 2026-09-27: a desktop GB10 Grace Blackwell system, popular for running large models locally because of its 128 GB of unified memory. Sources, all NVIDIA's DGX Spark page (nvidia.com/en-us/products/workstations/dgx-spark/): 128 GB LPDDR5x \"coherent unified system memory\", 256-bit, 273 GB/s; \"Up to 1 PFLOP FP4\", footnoted as theoretical FP4 with sparsity, so fp4TFLOPS here is the dense half, 500; GB10 TDP 140 W, footnoted as the whole chip (CPU and GPU); 240 W power supply. Compute capability 12.1 from NVIDIA's TensorRT support matrix (\"12.1 (DGX Spark)\"). vramGB is the full unified memory; the CPU shares it, so less is available to models in practice. The CPU is Arm (20-core, 10 Cortex-X925 + 10 Cortex-A725), so software needs aarch64 (Linux arm64) builds: PyTorch's CUDA 13 aarch64 wheels include Blackwell code, and its RELEASE.md lists 12.0 for aarch64 builds, which covers 12.1.\n"
    },
    {
      "id": "nvidia-gtx-1080-ti",
      "vendor": "nvidia",
      "name": "GTX 1080 Ti",
      "category": "consumer",
      "architecture": "Pascal",
      "releaseYear": 2017,
      "launchDate": "Mar 10, 2017 (announced Feb 28 at GDC)",
      "targetId": "6.1",
      "vramGB": 11,
      "memoryType": "GDDR5X",
      "memoryBandwidthGBs": 484,
      "fp32TFLOPS": 11.3,
      "transistorCountBillion": 12,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 28,
      "shaderCoreCount": 3584,
      "peakClockMHz": 1582,
      "tdpWatts": 250,
      "pcieGen": "PCIe Gen3 x16",
      "msrpUSD": 699,
      "notes": "Added 2026-09 — the user asked for older Titan/GTX cards specifically; GTX 1080 Ti was one of the most widely adopted early deep-learning GPUs thanks to strong FP32 throughput per dollar, predating this collection's previous oldest consumer-tier Pascal-class part. Core facts (3,584 CUDA cores, 11GB GDDR5X at 11Gbps, 12 billion transistors, $699 launch price, March 10 2017 availability) are from NVIDIA's own official press release (nvidianews.nvidia.com, \"NVIDIA Introduces the Beastly GeForce GTX 1080 Ti\", Feb 28 2017), read via pdftotext.\nmemoryBandwidthGBs (484), peakClockMHz (1582 boost), tdpWatts (250), and the 352-bit memory bus width (not itself stored as a field) are not in that press release's own text but are the same figures NVIDIA distributed on its press spec sheet at the same launch event, reported consistently by outlets present at the briefing (e.g. GamersNexus's \"Official nVidia GTX 1080 Ti Specs\" writeup) and corroborated unanimously across independent GPU databases — treated as reliable despite not being quoted verbatim in the press-release PDF itself. fp32TFLOPS (11.3) is derived from those same clock/core figures using this site's standard vector-throughput formula (cores × boost clock × 2), not independently published as a headline TFLOPS number by NVIDIA for this SKU. No Tensor/ matrix-throughput fields — Pascal predates Tensor Cores (introduced with Volta the following year), and GP102 has no dedicated FP16 Matrix path (fp16TFLOPSVector also omitted — not stated at 2x rate for this gaming-tier Pascal die, unlike Tesla P100's GP100).\n2026-09-15 — added computeUnitCount (28 SMs). NVIDIA publishes the 3,584 CUDA-core total (already in shaderCoreCount) but not the SM count; 28 follows from Pascal's fixed 128 FP32 cores per SM on the consumer GP10x dies (3,584 / 128 = 28) — GP102 with 28 of its 30 SMs enabled. matrixCoreCount stays correctly empty: no Tensor Cores on Pascal.\n"
    },
    {
      "id": "nvidia-h100-sxm5-80gb",
      "vendor": "nvidia",
      "name": "H100 SXM5 80GB",
      "category": "datacenter",
      "architecture": "Hopper",
      "releaseYear": 2022,
      "targetId": "9.0",
      "vramGB": 80,
      "memoryType": "HBM3",
      "memoryBandwidthGBs": 3350,
      "fp64TFLOPS": 34,
      "fp32TFLOPS": 67,
      "fp64TFLOPSMatrix": 67,
      "tf32TFLOPS": 494.5,
      "bf16TFLOPS": 989.5,
      "fp16TFLOPS": 989.5,
      "fp8TFLOPS": 1979,
      "int8TOPS": 1979,
      "fp16TFLOPSSparse": 1979,
      "fp8TFLOPSSparse": 3958,
      "int8TOPSSparse": 3958,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 132,
      "shaderCoreCount": 16896,
      "matrixCoreCount": 528,
      "tdpWatts": 700,
      "formFactor": "SXM",
      "pcieGen": "PCIe Gen5",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 900,
      "multiInstanceSupport": "Up to 7 MIGs @ 10GB each",
      "notes": "All figures sourced from NVIDIA's official H100 product page (nvidia.com/en-us/data-center/h100/, \"Product Specifications\" table, H100 SXM column), pulled 2026-09-01.\ntf32TFLOPS, bf16TFLOPS, fp16TFLOPS, fp8TFLOPS, and int8TOPS are dense (non-sparse) figures, derived by halving NVIDIA's own published \"with sparsity\" headline numbers (989/1,979/1,979/3,958/3,958 respectively) — the page's own footnote (\"* With sparsity\") confirms this 2x convention, applied consistently across this site's Ampere/Hopper/Ada/Blackwell entries. The ...Sparse fields hold the vendor's directly published headline figures. fp64TFLOPS (34) and fp64TFLOPSMatrix (67) are NOT marked with the sparsity footnote in NVIDIA's table, so both are used as-is with no halving.\ninterconnectBandwidthGBs (900) is NVLink; the SXM module also exposes a PCIe Gen5 x16 host link at 128GB/s (pcieGen records the generation only). tdpWatts (700) is \"Up to 700W (configurable)\" per NVIDIA — the actual board default may run lower. No transistor count, L2 cache size, or process node published on this product page or in the quick-reference datasheet (the SM / CUDA-core / Tensor-core counts were missing for the same reason and have since been added from NVIDIA's Hopper architecture blog — see below). No msrpUSD — H100 ships through OEM/server partners, not direct retail.\n2026-09-15 — added computeUnitCount (132 SMs), shaderCoreCount (16,896 FP32 CUDA cores) and matrixCoreCount (528 fourth-generation Tensor Cores) from NVIDIA's own \"NVIDIA Hopper Architecture In-Depth\" developer blog post (developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/), which gives the shipping H100 SXM5 configuration explicitly rather than the full GH100 die (144 SMs / 18,432 FP32 cores / 576 Tensor Cores — 12 SMs are fused off on every SXM5 part). None of these three figures appear on the H100 product page or in the quick-reference datasheet used for every other field here.\n"
    },
    {
      "id": "nvidia-h200-sxm-141gb",
      "vendor": "nvidia",
      "name": "H200 SXM 141GB",
      "category": "datacenter",
      "architecture": "Hopper",
      "releaseYear": 2024,
      "targetId": "9.0",
      "vramGB": 141,
      "memoryType": "HBM3e",
      "memoryBandwidthGBs": 4800,
      "fp64TFLOPS": 34,
      "fp32TFLOPS": 67,
      "fp64TFLOPSMatrix": 67,
      "tf32TFLOPS": 494.5,
      "bf16TFLOPS": 989.5,
      "fp16TFLOPS": 989.5,
      "fp8TFLOPS": 1979,
      "int8TOPS": 1979,
      "fp16TFLOPSSparse": 1979,
      "fp8TFLOPSSparse": 3958,
      "int8TOPSSparse": 3958,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 132,
      "shaderCoreCount": 16896,
      "matrixCoreCount": 528,
      "tdpWatts": 700,
      "formFactor": "SXM",
      "pcieGen": "PCIe Gen5",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 900,
      "multiInstanceSupport": "Up to 7 MIGs @ 18GB each",
      "notes": "Same Hopper compute die as H100 — H200 is a memory upgrade (more capacity, higher bandwidth), not a compute upgrade; compute figures are identical to the H100 SXM entry. Source: NVIDIA's official H200 product page (nvidia.com/en-us/data-center/h200/, \"Specifications\" table, H200 SXM column), pulled 2026-09-01. NVIDIA marks this table \"Preliminary specifications. May be subject to change.\"\ntf32TFLOPS, bf16TFLOPS, fp16TFLOPS, fp8TFLOPS, and int8TOPS are dense (non-sparse) figures, derived by halving NVIDIA's published \"with sparsity\" headline numbers (989/1,979/1,979/3,958/3,958), per the page's own footnote — same methodology as the H100 entry. The ...Sparse fields hold the vendor's directly published headline (sparse) figures. fp64 rows are not marked with the sparsity footnote, so fp64TFLOPS (34) and fp64TFLOPSMatrix (67) are used as-is.\nmultiInstanceSupport partition size (18GB) differs from H100's (10GB) to reflect H200's larger 141GB VRAM pool split across the same 7 MIG slices. No transistor count, L2 cache, or process node published on this page (SM / CUDA-core / Tensor-core counts have since been added from H200's shared GH100 die — see below). No msrpUSD — ships through OEM/server partners.\n2026-09-15 — added computeUnitCount (132 SMs), shaderCoreCount (16,896 FP32 CUDA cores) and matrixCoreCount (528 fourth-generation Tensor Cores). NVIDIA does not publish these for H200 specifically: they are H100's figures, carried over because H200 reuses the same GH100 compute die unchanged (same 4N die, same 132 of 144 SMs enabled) and changes only the memory stack — HBM3e instead of HBM3, 141GB instead of 80GB. That is exactly why every compute figure in this file already matches this collection's H100 SXM5 entry line for line, with only the memory and MIG-partition rows differing. Recorded as a same-die inference, not a vendor-stated figure; if NVIDIA ever publishes a differing count for H200, these three fields are the ones to revisit.\n"
    },
    {
      "id": "nvidia-k80",
      "vendor": "nvidia",
      "name": "K80",
      "category": "datacenter",
      "architecture": "Kepler",
      "releaseYear": 2014,
      "launchDate": "Nov 2014 (GA); board spec revised Jan 2015",
      "targetId": "3.7",
      "vramGB": 24,
      "memoryType": "GDDR5",
      "memoryBandwidthGBs": 480,
      "eccSupport": true,
      "fp64TFLOPS": 2.91,
      "fp32TFLOPS": 8.74,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 26,
      "shaderCoreCount": 4992,
      "tdpWatts": 300,
      "formFactor": "Dual-GPU PCIe, Full Height",
      "pcieGen": "PCIe Gen3",
      "notes": "Added 2026-09 alongside M40 and older AMD Radeon Instinct cards — K80 predates this site's previous oldest NVIDIA entry (2017's V100) by three years and, despite its age, is still occasionally encountered on legacy training infrastructure. Board-level figures (FP64/FP32 TFLOPS with GPU Boost, memory, bandwidth, power) sourced from NVIDIA's official \"Tesla K80\" overview PDF (nvidia.com/content/dam/en-zz/Solutions/Data-Center/ tesla-product-literature/nvidia-tesla-k80-overview.pdf); mechanical/power details (form factor, ECC, ~300W max board power) cross-checked against NVIDIA's official Board Specification document (BD-07317-001_v05, Jan 2015, nvidia.com/content/dam/en-zz/Solutions/Data-Center/ tesla-product-literature/Tesla-K80-BoardSpec-07317-001-v05.pdf). Both read via pdftotext.\nK80 is a genuinely dual-GPU board — two Tesla GK210 dies connected via an onboard PLX PCIe switch, each with 12GB GDDR5 (24GB total) and 2,496 CUDA cores (4,992 combined, the figure in shaderCoreCount here) — unlike every other entry in this collection, which is a single die. NVIDIA sells and benchmarks the K80 as one board-level SKU, so fp64TFLOPS (2.91) and fp32TFLOPS (8.74) are the combined-board figures \"with NVIDIA GPU Boost\" from the overview PDF, not per-GPU. There is no separate host-facing NVLink/interconnect field populated — the board's only external interface is the single PCIe Gen3 x16 link shared by both dies; the PLX switch routing between them isn't modeled as an `interconnect` value.\ntargetId (\"3.7\", GK210's actual compute capability) is set for correctness. As of 2026-09 this now cross-links to the 10.0 and 10.1 CUDA entries (both list \"3.7\"), which predate this site's Maxwell-only 10.2+ entries — Kepler support was deprecated in 10.2 and dropped entirely in 11.0, so any current-day CUDA Toolkit genuinely cannot target Kepler. No process node, transistor count, or MSRP found in either sourced document — K80 shipped through OEM/server partners, not direct retail.\n2026-09-15 — added computeUnitCount (26 SMX blocks). NVIDIA's K80 datasheet publishes only the 4,992 CUDA-core board total; 26 follows from Kepler's fixed 192 FP32 cores per SMX (4,992 / 192 = 26) and matches the board configuration reported at launch — two GK210 dies with 13 of each die's 15 SMX blocks enabled. Kepler calls the block an SMX; computeUnitLabel says \"Streaming Multiprocessors\" for consistency with the rest of this collection's NVIDIA entries. As with every other figure in this file, 26 is the whole-board number: software sees two 13-SMX devices, not one 26-SMX device.\n"
    },
    {
      "id": "nvidia-l40s",
      "vendor": "nvidia",
      "name": "L40S",
      "category": "datacenter",
      "architecture": "Ada Lovelace",
      "releaseYear": 2023,
      "targetId": "8.9",
      "vramGB": 48,
      "memoryType": "GDDR6 (ECC)",
      "memoryBandwidthGBs": 864,
      "fp32TFLOPS": 91.6,
      "tf32TFLOPS": 183,
      "bf16TFLOPS": 362.05,
      "fp16TFLOPS": 362.05,
      "fp8TFLOPS": 733,
      "fp16TFLOPSSparse": 733,
      "fp8TFLOPSSparse": 1466,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 142,
      "shaderCoreCount": 18176,
      "matrixCoreCount": 568,
      "tdpWatts": 350,
      "formFactor": "PCIe dual-slot",
      "pcieGen": "PCIe Gen4",
      "notes": "fp16TFLOPS/fp8TFLOPS dense figures and their Sparse counterparts are directly published dense/sparse pairs from NVIDIA's official L40S datasheet (362.05/733 dense; 733/1466 with structured sparsity) — original source verified in an earlier session. fp32TFLOPS (91.6) and tf32TFLOPS (183 dense, 366 with sparsity) added 2026-09-01 from NVIDIA's current L40S product page (nvidia.com/en-us/data-center/l40s/, \"Specifications\" panel) via screenshot (the page renders this panel in a way get_page_text and WebFetch on the linked datasheet PDF both failed to extract as text — read visually instead). bf16TFLOPS assumed equal to fp16TFLOPS (NVIDIA doesn't list BF16 separately on this page; Ada Lovelace Tensor Cores run BF16 and FP16 at the same rate on every other entry in this collection where both are independently published).\nNo interconnect/interconnectBandwidthGBs — L40S has no NVLink, PCIe Gen4 x16 only. No transistor count or L2 cache found in text form (the page also shows an \"RT Core Performance: 212 TFLOPS\" ray-tracing figure, not included — out of scope for this AI/ML-focused schema). No public MSRP — L40S ships through the OEM/server-partner channel, not direct retail.\n2026-09-15 — added shaderCoreCount (18,176 CUDA cores) and matrixCoreCount (568 fourth-generation Tensor Cores), both read from the specifications table on NVIDIA's own L40S product page (nvidia.com/en-us/data-center/l40s/) — they were on the page all along and simply missed on the earlier pass, not newly published. Also added computeUnitCount (142 SMs), which NVIDIA does not state directly but which both published counts independently imply on Ada Lovelace's fixed per-SM structure: 18,176 / 128 FP32 cores per SM = 142, and 568 / 4 Tensor Cores per SM = 142. The same table's \"142 third-generation RT Cores\" is a third confirmation, since Ada carries exactly one RT Core per SM. Three agreeing figures is why this is recorded as a value rather than left blank.\n"
    },
    {
      "id": "nvidia-m40-24gb",
      "vendor": "nvidia",
      "name": "M40 24GB",
      "category": "datacenter",
      "architecture": "Maxwell",
      "releaseYear": 2015,
      "launchDate": "Nov 2015 (12GB, datasheet dated Jan16); 24GB variant datasheet dated Mar16",
      "targetId": "5.2",
      "vramGB": 24,
      "memoryType": "GDDR5",
      "memoryBandwidthGBs": 288,
      "fp64TFLOPS": 0.2,
      "fp32TFLOPS": 7,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 24,
      "shaderCoreCount": 3072,
      "tdpWatts": 250,
      "formFactor": "PCIe Full Height, Dual Slot",
      "pcieGen": "PCIe Gen3",
      "notes": "Added 2026-09 alongside K80 and older AMD Radeon Instinct cards to extend NVIDIA coverage further back — M40 was a widely-deployed early deep learning training GPU (Caffe/Torch era) and predates this site's previous oldest NVIDIA entry (2017's V100) by two years. All figures sourced directly from NVIDIA's official Tesla M40 datasheet, 24GB variant (images.nvidia.com/content/tesla/pdf/78071_Tesla_M40_24GB_Print_Datasheet_LR.PDF, footer dated MAR16), read via pdftotext. NVIDIA also published an otherwise-identical 12GB datasheet (nvidia-teslam40-datasheet.pdf, dated Jan16) with the same architecture/CUDA-core-count/TFLOPS/bandwidth/power figures — this entry represents the 24GB SKU, matching this site's general preference for the higher-memory variant where both exist (see V100 SXM2 32GB's notes for the same reasoning).\nfp32TFLOPS (7) is explicitly \"with NVIDIA GPU Boost\" per the datasheet's own Features list. eccSupport is intentionally omitted (not false) — unlike this site's other Tesla-era entries (K80, P100, V100), M40's official spec table has no ECC row at all, and Maxwell GM200's GDDR5 (rather than HBM2) memory controller doesn't get the same blanket ECC treatment those datasheets call out, so leaving it unspecified is more accurate than guessing either way. No Tensor/matrix-throughput fields — Maxwell predates Tensor Cores entirely (introduced with Volta), and this datasheet gives only a single FP32 rate, not a separate FP16 vector figure. No transistor count, process node, or MSRP — M40 shipped through OEM/server partners, not direct retail.\n2026-09-15 — added computeUnitCount (24 SMs) from Table 1 of NVIDIA's Volta architecture whitepaper (WP-08608-001_v1.1), which lists GM200 (Maxwell) at 24 SMs, 128 FP32 cores per SM and 3,072 FP32 cores per GPU — the same 3,072 already in shaderCoreCount here. Maxwell's own name for the block is SMM (as Kepler's is SMX); computeUnitLabel says \"Streaming Multiprocessors\" for consistency with the rest of this collection's NVIDIA entries. matrixCoreCount stays empty — no Tensor Cores on Maxwell.\n"
    },
    {
      "id": "nvidia-p100-sxm2-16gb",
      "vendor": "nvidia",
      "name": "P100 SXM2 16GB",
      "category": "datacenter",
      "architecture": "Pascal",
      "releaseYear": 2016,
      "launchDate": "Apr 2016 (announced at GTC); datasheet dated Oct 2016",
      "targetId": "6.0",
      "vramGB": 16,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 732,
      "eccSupport": true,
      "fp64TFLOPS": 5.3,
      "fp32TFLOPS": 10.6,
      "fp16TFLOPSVector": 21.2,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 56,
      "shaderCoreCount": 3584,
      "tdpWatts": 300,
      "formFactor": "SXM2",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 160,
      "notes": "Added 2026-09 alongside V100/T4/RTX 3090 to extend NVIDIA coverage further back — P100 predates this site's previous oldest NVIDIA entry (2020's A100) by four years and is still encountered on older training clusters. Core figures (CUDA core count, FP64/FP32/FP16 TFLOPS, memory size/bandwidth, TDP, ECC, form factor) sourced directly from NVIDIA's official Tesla P100 datasheet (nvidia.com/content/dam/en-zz/Solutions/ Data-Center/tesla-p100/pdf/nvidia-tesla-p100-datasheet.pdf, SXM2 column, footer dated Oct16), read via pdftotext. interconnectBandwidthGBs (160, NVLink) is from the companion Technical Overview PDF on the same page (nvidia-teslap100-techoverview.pdf), which explicitly states \"four NVLink connections per GPU — each delivering 40 GB/sec bi-directional interconnect bandwidth, Tesla P100 delivers 160 GB/s bidirectional bandwidth in total.\"\nfp16TFLOPSVector (21.2) is Pascal's native packed-FP16 throughput via CUDA cores, listed in the datasheet as \"Half-Precision Performance\" — correctly placed in this schema's vector (non-Tensor) field, not fp16TFLOPS/matrix, because Pascal has no Tensor Cores at all (introduced with Volta the following generation); fp16TFLOPS, bf16TFLOPS, tf32TFLOPS, fp8TFLOPS, and int8TOPS are all correctly omitted rather than left as zero. eccSupport is \"Yes\" per the datasheet's own note: \"Native support with no capacity or performance overhead\" (HBM2's on-package ECC, unlike GDDR5-based Kepler-generation cards which took a bandwidth/capacity hit for ECC). No transistor count, process node, PCIe generation, or MSRP in either sourced document — P100 shipped through OEM/server partners, not direct retail; a 12GB HBM2 SKU also exists but isn't modeled here.\n2026-09-15 — added computeUnitCount (56 SMs) from Table 1 of NVIDIA's Volta architecture whitepaper (WP-08608-001_v1.1), whose GPU-comparison table covers Tesla K40/M40/P100/V100 side by side and lists GP100 (Pascal) at 56 SMs, 64 FP32 cores per SM, 3,584 FP32 cores per GPU — that last figure matching the shaderCoreCount already recorded here from the P100 datasheet itself. matrixCoreCount stays correctly empty: Pascal has no Tensor Cores (introduced with Volta the following year).\n"
    },
    {
      "id": "nvidia-rtx-3060-12gb",
      "vendor": "nvidia",
      "name": "RTX 3060 12GB",
      "category": "consumer",
      "architecture": "Ampere",
      "releaseYear": 2021,
      "targetId": "8.6",
      "vramGB": 12,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 360,
      "fp32TFLOPS": 12.8,
      "shaderCoreCount": 3584,
      "peakClockMHz": 1780,
      "tdpWatts": 170,
      "pcieGen": "PCIe Gen4",
      "msrpUSD": 329,
      "notes": "Added 2026-09-27 as one of the mid-range cards local-LLM users actually buy (12 GB for $329 at launch). Sources, all NVIDIA: 3,584 CUDA cores, 1.78 GHz boost, 12 GB GDDR6 and 170 W Graphics Card Power from the RTX 3060 family product page (nvidia.com/en-us/geforce/graphics-cards/30-series/ rtx-3060-3060ti/); 360 GB/s from NVIDIA's own launch Q&A (\"GeForce RTX 3060 is 192-bit and GDDR6 (at 15 Gbps), delivering 360GB/s of bandwidth\", nvidia.com/en-us/geforce/news/game-on-you-asked-we-answered-qa/); $329 and the late-February 2021 availability from NVIDIA's announcement article (nvidia.com/en-us/geforce/news/geforce-rtx-3060/). compute capability 8.6 from NVIDIA's GeForce compare page (CUDA Capability row, RTX 30 Series). fp32TFLOPS is derived, not published: 3,584 cores × 2 FLOPs × 1.78 GHz = 12.76. NVIDIA publishes no Tensor-core throughput for this card on its product page, so those fields are omitted rather than filled from secondary sources. The 8 GB RTX 3060 is a different, narrower-bus card and is not this entry.\n"
    },
    {
      "id": "nvidia-rtx-3090",
      "vendor": "nvidia",
      "name": "RTX 3090",
      "category": "consumer",
      "architecture": "Ampere",
      "releaseYear": 2020,
      "launchDate": "Sep 24, 2020",
      "targetId": "8.6",
      "vramGB": 24,
      "memoryType": "GDDR6X",
      "memoryBandwidthGBs": 936,
      "cacheMB": 6,
      "fp32TFLOPS": 35.6,
      "tf32TFLOPS": 35.6,
      "bf16TFLOPS": 71.2,
      "fp16TFLOPS": 71.2,
      "int8TOPS": 284.7,
      "fp16TFLOPSSparse": 142.4,
      "int8TOPSSparse": 569.4,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 82,
      "shaderCoreCount": 10496,
      "matrixCoreCount": 328,
      "tdpWatts": 350,
      "pcieGen": "PCIe Gen4",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 112.5,
      "msrpUSD": 1499,
      "notes": "Added 2026-09 alongside V100/P100/T4 — RTX 3090 was this site's own README-flagged example of a missing older-but-still-common AI/ML GPU (it remains popular for local LLM inference/fine-tuning due to its 24GB VRAM at consumer pricing). Core compute figures, CUDA/Tensor/SM counts, memory config, L2 cache, and TGP are all read directly from NVIDIA's own \"NVIDIA RTX Blackwell GPU Architecture\" whitepaper (images.nvidia.com/ aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf, Appendix A, Table 3: \"GeForce RTX 5090 vs GeForce RTX 4090 vs GeForce RTX 3090 Specs\" — the same table this site's RTX 4090 entry already cites for its own figures, extended here to its RTX 3090 column), read via pdftotext, 2026-09. Board-level fields (boost clock, memory config, PCIe gen, TGP) cross-checked against nvidia.com's own current RTX 3090 product spec page and matched exactly.\nfp32TFLOPS/tf32TFLOPS (35.6) match because Ampere's non-Tensor FP32 and dense TF32-Tensor rate are numerically identical on this die, same pattern already documented on this site's RTX 4090 entry; fp16TFLOPSVector is omitted for the same reason (whitepaper's \"Peak FP16 TFLOPS (non-Tensor)\" is also 35.6 — redundant with fp32TFLOPS). bf16TFLOPS/ fp16TFLOPS (71.2 dense, 142.4 sparse) use the whitepaper's \"...with FP32 Accumulate\" row, matching this site's established convention (see the RTX 4090 entry's notes) rather than the higher \"...with FP16 Accumulate\" row (142.3/284.6) also present in the same table. fp8TFLOPS/fp4TFLOPS are correctly omitted, not zero — the whitepaper lists both as \"N/A\" for RTX 3090 (FP8 Tensor arrived with Ada, FP4 with Blackwell). int4TOPS is likewise not published for this generation.\ninterconnectBandwidthGBs (112.5, NVLink) is a well-corroborated secondary-source figure (multiple independent reviewer/community sources), not read off an NVIDIA table directly — RTX 3090 is one of the few GeForce cards with an NVLink bridge connector (dropped starting with 3090 Ti), and NVIDIA's own marketing material confirms NVLink support without stating the exact GB/s figure. msrpUSD (1499) is the original Sept 2020 Founders Edition launch price, widely corroborated across contemporaneous outlets rather than quoted verbatim from NVIDIA's own launch page (which confirms the launch but not the price in the specific text pulled). No transistor count or process node — not in the sourced whitepaper table (publicly reported elsewhere as Samsung 8N, but not confirmed here directly, so omitted rather than assumed).\n"
    },
    {
      "id": "nvidia-rtx-4060-ti-16gb",
      "vendor": "nvidia",
      "name": "RTX 4060 Ti 16GB",
      "category": "consumer",
      "architecture": "Ada Lovelace",
      "releaseYear": 2023,
      "targetId": "8.9",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 288,
      "cacheMB": 32,
      "fp32TFLOPS": 22.1,
      "int8TOPSSparse": 353,
      "shaderCoreCount": 4352,
      "peakClockMHz": 2540,
      "tdpWatts": 165,
      "pcieGen": "PCIe Gen4",
      "msrpUSD": 499,
      "notes": "Added 2026-09-27: the cheapest 16 GB NVIDIA card of its generation, a common local-LLM pick. Sources, all NVIDIA: 4,352 CUDA cores, 2.54 GHz boost, \"22 TFLOPS\" shader and \"353 AI TOPS\" from the RTX 4060 family product page (nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4060-4060ti/), whose Total Graphics Power row reads \"165 or 160\" for the 16 GB / 8 GB models, in the same order as its memory row. 288 GB/s from NVIDIA's VRAM deep-dive (nvidia.com/en-us/geforce/news/rtx-40-series-vram-video-memory-explained/), which describes the RTX 4060 Ti and its 32 MB L2 as \"an Ada GPU with 288 GB/sec of peak memory bandwidth\". The launch article says the 16 GB model has \"additional graphics memory but otherwise identical specifications\" and starts at $499. fp32TFLOPS 22.1 is cores × 2 × 2.54 GHz, matching NVIDIA's rounded 22. int8TOPSSparse holds the \"AI TOPS\" figure, which on Ada is the FP8/INT8 rate with sparsity (the same reading as the RTX 4090 entry's 1,321). The card's PCIe interface is x8 electrically.\n"
    },
    {
      "id": "nvidia-rtx-4090",
      "vendor": "nvidia",
      "name": "RTX 4090",
      "category": "consumer",
      "architecture": "Ada Lovelace",
      "releaseYear": 2022,
      "targetId": "8.9",
      "vramGB": 24,
      "memoryType": "GDDR6X",
      "memoryBandwidthGBs": 1008,
      "cacheMB": 72,
      "fp32TFLOPS": 82.6,
      "tf32TFLOPS": 82.6,
      "bf16TFLOPS": 165.2,
      "fp16TFLOPS": 165.2,
      "fp8TFLOPS": 330.3,
      "int8TOPS": 660.6,
      "fp16TFLOPSSparse": 330.4,
      "fp8TFLOPSSparse": 660.6,
      "int8TOPSSparse": 1321.2,
      "processNode": "TSMC 4N",
      "transistorCountBillion": 76.3,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 128,
      "shaderCoreCount": 16384,
      "matrixCoreCount": 512,
      "peakClockMHz": 2520,
      "tdpWatts": 450,
      "formFactor": "PCIe 3-slot",
      "pcieGen": "PCIe Gen4",
      "msrpUSD": 1599,
      "notes": "Rewritten 2026-09-01 using NVIDIA's official \"NVIDIA RTX Blackwell GPU Architecture\" whitepaper (images.nvidia.com/aem-dam/Solutions/geforce/ blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf), which includes a side-by-side RTX 3090/4090/5090 spec table (Table 1) plus a fuller Appendix A table — read via pdftotext after WebFetch failed on this PDF's binary structure. This single source super-supersedes the earlier, more scattered pull and confirms every previously-recorded figure was already correct.\ntf32TFLOPS, bf16TFLOPS, fp16TFLOPS, and fp8TFLOPS are all dense Tensor-Core figures using FP32 accumulate (NVIDIA's own convention, matching this site's established methodology); the ...Sparse fields and int8TOPSSparse hold the structured-sparsity (\"with sparsity\") figures from the same rows. fp32TFLOPS (82.6) is the non-Tensor/vector \"Peak FP32 TFLOPS\" row — NVIDIA's whitepaper shows FP32, FP16, and BF16 non-Tensor throughput as numerically identical (82.6) on Ada's unified shader cores; tf32TFLOPS dense is also 82.6 for the same underlying reason (TF32 Tensor dense = FP32 vector rate on this die). int8TOPS/int8TOPSSparse (660.6/1321.2) match the card's own marketing \"1,321 AI TOPS\" figure exactly, confirming that headline number is the FP8(-with-FP16-accumulate)/INT8 sparse rate, not the FP32-accumulate figures used elsewhere in this file. NVIDIA does not list an FP4 Tensor row for RTX 4090 (Blackwell-only feature) — fp4TFLOPS omitted, not zero.\nprocessNode (\"TSMC 4N\"), transistorCountBillion (76.3), cacheMB (72, from \"L2 Cache Size: 73,728 KB\"), computeUnitCount (128 SMs), matrixCoreCount (512 Tensor Cores, 4th Gen), and shaderCoreCount/ peakClockMHz (confirmed unchanged: 16,384 CUDA cores, 2,520 MHz boost) are all read directly from the same whitepaper table. Die size (608.5 mm²) and RT Core count/TFLOPS are published but out of scope for this AI/ML-focused schema.\n"
    },
    {
      "id": "nvidia-rtx-5060-ti-16gb",
      "vendor": "nvidia",
      "name": "RTX 5060 Ti 16GB",
      "category": "consumer",
      "architecture": "Blackwell",
      "releaseYear": 2025,
      "targetId": "12.0",
      "vramGB": 16,
      "memoryType": "GDDR7",
      "memoryBandwidthGBs": 448,
      "fp32TFLOPS": 23.7,
      "fp4TFLOPS": 379.5,
      "shaderCoreCount": 4608,
      "peakClockMHz": 2570,
      "tdpWatts": 180,
      "pcieGen": "PCIe Gen5",
      "notes": "Added 2026-09-27: the cheapest 16 GB Blackwell card, with native FP4. Sources, all NVIDIA: 4,608 CUDA cores, 2.57 GHz boost, \"759 AI TOPS\", \"16 GB / 8 GB GDDR7\", 128-bit and 448 GB/sec from the RTX 50 Series table on NVIDIA's GeForce compare page (nvidia.com/en-us/geforce/graphics-cards/compare/), which also confirms this site's RTX 5090 bandwidth (1,792 GB/sec). 180 W Total Graphics Power from the RTX 5060 family product page. fp32TFLOPS 23.7 is cores × 2 × 2.57 GHz. fp4TFLOPS 379.5 is half the 759 \"AI TOPS\", which on Blackwell is the FP4 rate with sparsity (the same reading as the RTX 5090 entry's 3,352 / 1,676). No MSRP recorded: not found on an NVIDIA page.\n"
    },
    {
      "id": "nvidia-rtx-5090",
      "vendor": "nvidia",
      "name": "RTX 5090",
      "category": "consumer",
      "architecture": "Blackwell",
      "releaseYear": 2025,
      "targetId": "12.0",
      "vramGB": 32,
      "memoryType": "GDDR7",
      "memoryBandwidthGBs": 1792,
      "cacheMB": 96,
      "fp32TFLOPS": 104.8,
      "tf32TFLOPS": 104.8,
      "bf16TFLOPS": 209.5,
      "fp16TFLOPS": 209.5,
      "fp8TFLOPS": 419,
      "fp4TFLOPS": 1676,
      "int8TOPS": 838,
      "fp16TFLOPSSparse": 419,
      "fp8TFLOPSSparse": 838,
      "int8TOPSSparse": 1676,
      "processNode": "TSMC 4N",
      "transistorCountBillion": 92.2,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 170,
      "shaderCoreCount": 21760,
      "matrixCoreCount": 680,
      "peakClockMHz": 2407,
      "tdpWatts": 575,
      "formFactor": "PCIe dual-slot",
      "pcieGen": "PCIe Gen5",
      "msrpUSD": 1999,
      "notes": "Rewritten 2026-09-01 using NVIDIA's official \"NVIDIA RTX Blackwell GPU Architecture\" whitepaper (images.nvidia.com/aem-dam/Solutions/geforce/ blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf) — its Appendix A (\"Blackwell GB202 GPU\") table gives a full RTX 3090/4090/5090 side-by-side comparison, read via pdftotext after WebFetch failed on this PDF's binary structure. Same source and methodology as the RTX 4090 entry; fp16TFLOPS/fp8TFLOPS (209.5/419 dense, FP32-accumulate) were already correct from an earlier pull and are confirmed unchanged here.\ntf32TFLOPS, bf16TFLOPS, fp16TFLOPS, and fp8TFLOPS are dense Tensor-Core figures using FP32 accumulate; the ...Sparse fields and int8TOPSSparse hold the structured-sparsity figures from the same rows. fp32TFLOPS (104.8) is the non-Tensor/vector \"Peak FP32 TFLOPS\" row — identical to the non-Tensor FP16/BF16 rows on Blackwell's unified shader cores, and to tf32TFLOPS dense for the same reason as the RTX 4090 entry. fp4TFLOPS (1,676 dense; 3,352 with sparsity, no dedicated sparse field in this schema) is new to Blackwell — this is also the exact source of the card's marketing \"3,352 AI TOPS\" figure (FP4-with-sparsity), the same way RTX 4090's \"1,321 AI TOPS\" turned out to be its FP8/INT8 sparse rate. int8TOPS/int8TOPSSparse (838/1,676) match the FP8-with-FP16- accumulate row exactly (same rate on this die).\nprocessNode (\"TSMC 4N\"), transistorCountBillion (92.2 — RTX 5090 uses the full, uncut GB202 die's transistor budget even though only 170 of 192 SMs are enabled), cacheMB (96, from \"L2 Cache Size: 98,304 KB\"), computeUnitCount (170 SMs), matrixCoreCount (680 Tensor Cores, 5th Gen), shaderCoreCount (21,760 CUDA cores), and peakClockMHz (2,407 MHz boost) are all read directly from the same whitepaper table. Die size (750 mm²) and RT Core count/TFLOPS are published but out of scope for this AI/ML-focused schema.\n"
    },
    {
      "id": "nvidia-rtx-6000-ada-generation",
      "vendor": "nvidia",
      "name": "RTX 6000 Ada Generation",
      "category": "workstation",
      "architecture": "Ada Lovelace",
      "releaseYear": 2022,
      "targetId": "8.9",
      "vramGB": 48,
      "memoryType": "GDDR6 (ECC)",
      "memoryBandwidthGBs": 960,
      "fp32TFLOPS": 91.1,
      "bf16TFLOPS": 364.25,
      "fp16TFLOPS": 364.25,
      "fp8TFLOPS": 728.5,
      "fp16TFLOPSSparse": 728.5,
      "fp8TFLOPSSparse": 1457,
      "processNode": "TSMC 4N",
      "transistorCountBillion": 76.3,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 142,
      "shaderCoreCount": 18176,
      "matrixCoreCount": 568,
      "tdpWatts": 300,
      "formFactor": "PCIe dual-slot",
      "coolingType": "Active",
      "pcieGen": "PCIe Gen4",
      "msrpUSD": 6800,
      "notes": "fp32TFLOPS (91.1) confirmed 2026-09-01 directly from NVIDIA's current product page (nvidia.com/en-us/products/workstations/rtx-6000/, \"Single-Precision Performance: 91.1 TFLOPS\"). fp8TFLOPS/fp16TFLOPS remain derived, not directly published as separate dense figures: the same product page states \"Tensor Performance: 1,457 AI TOPS\" footnoted as \"Theoretical FP8 TOPS using the sparsity feature\" — halving gives dense FP8 (728.5, → fp8TFLOPSSparse: 1457), and halving again gives an estimated dense FP16 (364.25, → fp16TFLOPSSparse: 728.5), cross-checked against the L40S entry (same AD102 die, same 362.05/733 dense FP16/FP8 split) as a sanity check. bf16TFLOPS assumed equal to fp16TFLOPS — not separately published, same assumption made on the L40S entry for the same die family. No tf32TFLOPS — unlike the RTX 4090/5090 whitepaper, no source was found confirming TF32 dense equals the FP32 vector rate specifically for this SKU, so it's omitted rather than assumed.\nshaderCoreCount (18,176 CUDA cores), matrixCoreCount (568 4th-gen Tensor Cores), transistorCountBillion (76.3), and processNode (\"TSMC 4N\") come from a secondary source (an NVIDIA-authored datasheet mirrored by a cloud reseller, acecloud.ai) after NVIDIA's own gated datasheet link (resources.nvidia.com) and the mirror's own compute/TFLOPS table proved corrupted/unreliable on text extraction (looked like leftover Hopper template values, not RTX 6000 Ada's actual figures — discarded rather than used). The core counts, transistor count, and process are treated as reliable despite the secondary source because they're internally consistent (18,176 / 128 CUDA-cores-per-SM = 142 SMs exactly, used as computeUnitCount) and match AD102's well-established public transistor budget shared with the RTX 4090 entry. cacheMB and peakClockMHz omitted — not found in any source consulted.\nformFactor, coolingType, pcieGen, and tdpWatts (300W) are directly from NVIDIA's product page Specifications panel. msrpUSD ($6,800) remains an approximate street/channel price, not a datasheet-listed figure — RTX 6000 Ada sells through workstation partners rather than direct retail.\n"
    },
    {
      "id": "nvidia-rtx-pro-6000-blackwell",
      "vendor": "nvidia",
      "name": "RTX PRO 6000 Blackwell",
      "category": "workstation",
      "architecture": "Blackwell",
      "releaseYear": 2025,
      "launchDate": "2025-03-18 (announced; general availability April 2025)",
      "targetId": "12.0",
      "vramGB": 96,
      "memoryType": "GDDR7 (ECC)",
      "memoryBandwidthGBs": 1792,
      "eccSupport": true,
      "fp32TFLOPS": 125,
      "tf32TFLOPS": 125,
      "bf16TFLOPS": 250,
      "fp16TFLOPS": 250,
      "fp8TFLOPS": 500,
      "fp4TFLOPS": 2000,
      "int8TOPS": 1000,
      "fp16TFLOPSSparse": 500,
      "fp8TFLOPSSparse": 1000,
      "int8TOPSSparse": 2000,
      "processNode": "TSMC 4N",
      "transistorCountBillion": 92.2,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 188,
      "shaderCoreCount": 24064,
      "matrixCoreCount": 752,
      "tdpWatts": 600,
      "formFactor": "PCIe dual-slot",
      "coolingType": "Active (Double Flow Through)",
      "pcieGen": "PCIe Gen5",
      "notes": "Added 2026-09-06 as part of a Blackwell/RTX 50-series freshness check: this site already tracked the previous-generation workstation flagship (RTX 6000 Ada Generation) but not its current-generation successor.\nOnly two figures are directly published on NVIDIA's own product page (nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/): fp32TFLOPS (125, \"Single-Precision Performance\") and a single AI figure, \"4,000 TOPS,\" footnoted \"Theoretical FP4 TOPS using sparsity.\" Every other compute figure here is DERIVED from those two using the same precision-ladder relationship NVIDIA's own Blackwell architecture whitepaper establishes for the RTX 5090 (same GB202 die family, see that entry's notes): each step down in precision doubles dense throughput versus the step above (tf32 = fp32; bf16 = fp16 = 2x tf32; fp8 = 2x fp16; fp4 = 2x fp8), and each precision's sparse figure is 2x its own dense figure. Working forward from fp32=125 gives fp4 dense=2000 -> fp4 sparse=4000 -- which lands exactly on NVIDIA's own independently-published 4,000 TOPS figure, a strong cross-check that the assumed ratios hold for this die too, not just the 5090. int8TOPS/int8TOPSSparse are set equal to fp8TFLOPSSparse/2x that (1000/2000), mirroring the RTX 5090 entry where int8TOPS matches its FP8-sparse rate exactly.\ntransistorCountBillion (92.2) and processNode (\"TSMC 4N\") are carried over from the RTX 5090 entry, not independently re-published for this SKU -- both cards use the same physical GB202 die (this one effectively unlocked to 188 of the die's SMs, vs. 170 enabled on the 5090), and transistor count/process are fixed properties of the die itself, not of which SMs are enabled. computeUnitCount (188 SMs), shaderCoreCount (24,064 CUDA cores), and matrixCoreCount (752 5th-gen Tensor Cores) come from third-party reporting (Tom's Hardware/TechPowerUp/VideoCardz pricing coverage), not a spec sheet on NVIDIA's own product page (which lists no core counts at all) -- treated as reliable because they're internally consistent with each other (24,064 / 128 CUDA cores per SM = 188 SMs exactly; 752 / 188 = 4.0 Tensor Cores per SM, the same ratio the RTX 5090's whitepaper-sourced 680/170 = 4.0 figure shows).\nNo msrpUSD: NVIDIA never announced an official MSRP for this card -- early 2025 retailer/preorder listings for the boxed Workstation Edition ranged roughly $7,700-8,600, and NVIDIA has since raised its own Marketplace listing price twice without an announcement (to $13,250 in June 2026, then $16,000 in August 2026, per contemporaneous Tom's Hardware/VideoCardz/TechPowerUp coverage) -- picking any single number here would misrepresent an unusually volatile, vendor-unannounced pricing history as a stable list price.\nA distinct \"RTX PRO 6000 Blackwell Server Edition\" (passive cooling, for OEM servers) and a \"Max-Q\" variant (300W, blower-style cooler) also exist; this entry covers only the workstation/active-cooling edition described above, matching this collection's existing convention of one entry per primary SKU rather than every cooling/form-factor variant.\n"
    },
    {
      "id": "nvidia-t4",
      "vendor": "nvidia",
      "name": "T4",
      "category": "datacenter",
      "architecture": "Turing",
      "releaseYear": 2018,
      "launchDate": "Sep 2018 (announced); product brief dated Dec 2018",
      "targetId": "7.5",
      "vramGB": 16,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 300,
      "eccSupport": true,
      "fp32TFLOPS": 8.1,
      "fp16TFLOPS": 65,
      "int8TOPS": 130,
      "int4TOPS": 260,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 40,
      "shaderCoreCount": 2560,
      "matrixCoreCount": 320,
      "tdpWatts": 70,
      "formFactor": "Low-Profile PCIe",
      "pcieGen": "PCIe Gen3",
      "interconnect": "PCIe only (no NVLink)",
      "interconnectBandwidthGBs": 32,
      "notes": "Added 2026-09 alongside V100/P100/RTX 3090 — T4 is one of the most widely deployed inference GPUs in the field (cloud instances, edge inference) and was a notable gap in this site's NVIDIA coverage. All figures sourced directly from NVIDIA's official T4 datasheet (nvidia.com/content/dam/en-zz/Solutions/Data-Center/tesla-t4/ t4-tensor-core-datasheet-951643.pdf, footer dated Mar19), read via pdftotext.\nfp16TFLOPS (65) is the datasheet's \"Mixed-Precision (FP16/FP32)\" Tensor Core row — Turing's Tensor Cores (2nd gen) predate structured sparsity (introduced with Ampere), so there's no dense/sparse split and no fp16TFLOPSSparse value; likewise int8TOPS (130) and int4TOPS (260) are single dense figures with no sparse counterpart. No BF16, TF32, or FP8 Tensor rows exist for Turing (BF16/TF32 arrived with Ampere, FP8 with Hopper/Ada) — correctly omitted rather than left as zero. fp16TFLOPSVector is also omitted: the datasheet doesn't give a separate non-Tensor FP16 figure.\ninterconnect is PCIe-only — T4 has no NVLink, unlike the SXM-form-factor datacenter parts elsewhere in this collection; interconnectBandwidthGBs (32) is the x16 PCIe Gen3 host link bandwidth, explicitly labeled \"Interconnect Bandwidth: 32 GB/sec\" in the same table (separate from the \"System Interface: x16 PCIe Gen3\" row). tdpWatts (70) matches T4's signature low-profile single-slot design — no external power connector needed. No transistor count, process node, or MSRP in this datasheet — T4 shipped through OEM/server/cloud partners, not direct retail.\n2026-09-15 — added computeUnitCount (40 SMs). NVIDIA's T4 datasheet publishes 2,560 CUDA cores and 320 Tensor Cores (both already recorded here) but no SM count; 40 follows independently from either one on Turing's fixed per-SM structure — 2,560 / 64 FP32 cores per SM = 40, and 320 / 8 Tensor Cores per SM = 40. Two agreeing ratios, so it is recorded as a value rather than left blank.\n"
    },
    {
      "id": "nvidia-tesla-p40",
      "vendor": "nvidia",
      "name": "Tesla P40",
      "category": "datacenter",
      "architecture": "Pascal",
      "releaseYear": 2016,
      "launchDate": "Sep 2016 (datasheet footer dated Sep16)",
      "targetId": "6.1",
      "vramGB": 24,
      "memoryType": "GDDR5",
      "memoryBandwidthGBs": 346,
      "eccSupport": true,
      "fp32TFLOPS": 12,
      "int8TOPS": 47,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 30,
      "shaderCoreCount": 3840,
      "tdpWatts": 250,
      "formFactor": "Full-Height, Dual-Slot (4.4\" H x 10.5\" L)",
      "coolingType": "Passive",
      "pcieGen": "PCIe 3.0 x16",
      "notes": "Added 2026-09 — the most-discussed used budget card for local LLM inference (24GB for a fraction of an RTX 3090's price), and compute capability 6.1 means it lost CUDA support in 13.0 along with the rest of Pascal; CUDA 12.x is the last toolkit that can target it. Figures (architecture, 12 TFLOPS FP32, 47 TOPS INT8, 24GB, 346 GB/s, PCIe 3.0 x16, dual-slot full-height form factor, 250W max power, ECC) sourced directly from NVIDIA's official Tesla P40 datasheet (images.nvidia.com/ content/pdf/tesla/184427-Tesla-P40-Datasheet-NV-Final-Letter-Web.pdf, footer \"Sep16\"), read via pdftotext. NVIDIA marks both the FP32 and INT8 figures \"With Boost Clock Enabled\". The datasheet does not list a memory type; GDDR5 is the P40's documented memory (the P40 is a GP102 part, not the HBM2 GP100 used by the P100).\nshaderCoreCount (3840) and computeUnitCount (30 SMs) are the full GP102 die at 128 FP32 cores per SM — consistent with the datasheet's 12 TFLOPS: 3840 cores x 2 FLOP x ~1.53 GHz boost ≈ 11.8 TFLOPS. Not on the datasheet itself.\nfp16TFLOPSVector is deliberately omitted, and this matters more than any other number for this card: unlike the P100 (compute capability 6.0, fast packed FP16), GP102-based cards like the P40 execute FP16 at a tiny fraction of their FP32 rate. In practice, LLM inference on a P40 runs in FP32 or via integer-quantized kernels, not FP16. No Tensor Cores (introduced with Volta), so all matrix fields except int8TOPS are correctly omitted; int8TOPS is Pascal's DP4A instruction throughput on the CUDA cores, not Tensor-Core throughput.\n"
    },
    {
      "id": "nvidia-titan-rtx",
      "vendor": "nvidia",
      "name": "TITAN RTX",
      "category": "consumer",
      "architecture": "Turing",
      "releaseYear": 2018,
      "launchDate": "Dec 2018",
      "targetId": "7.5",
      "vramGB": 24,
      "memoryType": "GDDR6",
      "memoryBandwidthGBs": 672,
      "fp32TFLOPS": 16.3,
      "fp16TFLOPS": 130,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 72,
      "shaderCoreCount": 4608,
      "matrixCoreCount": 576,
      "peakClockMHz": 1770,
      "tdpWatts": 280,
      "pcieGen": "PCIe Gen3 x16",
      "interconnect": "NVLink (optional bridge)",
      "interconnectBandwidthGBs": 100,
      "msrpUSD": 2499,
      "notes": "Added 2026-09 alongside GTX 1080 Ti, TITAN X (Pascal), and TITAN V — the last of NVIDIA's consumer/prosumer TITAN line (no successor was released after Ampere), and notable for pairing Turing's Tensor Cores with 24GB of VRAM, unusually large for a non-datacenter card at the time. CUDA core count (4,608), Tensor core count (576), boost clock (1,770 MHz), memory config (24GB GDDR6), memory bandwidth (672 GB/s), Tensor performance (\"130 Tensor TFLOPs\"), and the optional NVLink bridge (100 GB/s, doubling effective VRAM to 48GB across two cards) are all quoted directly from NVIDIA's own official product page (nvidia.com/en-us/titan/titan-rtx/), read 2026-09.\nfp32TFLOPS (16.3) and tdpWatts (280) aren't on that specific page but are consistently corroborated across independent GPU databases and contemporaneous press coverage of the Dec 2018 launch. int8TOPS is intentionally omitted, not guessed — secondary sources disagree sharply on this figure (some report 130 TOPS, equal to the FP16 Tensor rate; others report roughly double), and no NVIDIA-published number was found to resolve the conflict. msrpUSD (2499) is the original launch price, well-corroborated but not itself quoted on the current product page (it no longer lists a price). RT Core count/TFLOPS are published but out of scope for this AI/ML-focused schema, matching this collection's existing convention (see RTX 4090's notes for the same treatment).\n2026-09-15 — added computeUnitCount (72 SMs). NVIDIA's TITAN RTX page gives 4,608 CUDA cores, 576 Tensor Cores and 72 RT Cores; on Turing's fixed per-SM structure all three independently give 72 SMs (4,608 / 64, 576 / 8, and one RT Core per SM). Recorded as a derived value rather than left blank because three published figures agree on it.\n"
    },
    {
      "id": "nvidia-titan-v",
      "vendor": "nvidia",
      "name": "TITAN V",
      "category": "consumer",
      "architecture": "Volta",
      "releaseYear": 2017,
      "launchDate": "Dec 7, 2017 (announced at NeurIPS/NIPS)",
      "targetId": "7.0",
      "vramGB": 12,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 653,
      "fp64TFLOPS": 7.4,
      "fp32TFLOPS": 14.9,
      "fp16TFLOPS": 110,
      "processNode": "TSMC 12nm FFN",
      "transistorCountBillion": 21.1,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 80,
      "shaderCoreCount": 5120,
      "matrixCoreCount": 640,
      "peakClockMHz": 1455,
      "tdpWatts": 250,
      "pcieGen": "PCIe Gen3 x16",
      "msrpUSD": 2999,
      "notes": "Added 2026-09 alongside GTX 1080 Ti, TITAN X (Pascal), and TITAN RTX — the first-ever consumer/prosumer card with Tensor Cores, using the same GV100 die as this collection's V100 entry. Confirmed directly from NVIDIA's own official press release (nvidianews.nvidia.com, \"NVIDIA TITAN V Transforms the PC into AI Supercomputer\", Dec 7 2017, read via pdftotext): 21.1 billion transistors, TSMC 12nm FFN process customized for NVIDIA, 12GB HBM2, 110 TFLOPS Tensor (deep learning) performance, and $2,999 launch price — NVIDIA never published a full traditional spec sheet for this card the way it does for GeForce/Tesla SKUs, so several fields below aren't from an NVIDIA-authored table.\nshaderCoreCount (5,120 CUDA cores), matrixCoreCount (640 Tensor Cores, 1st gen — identical counts to this collection's V100 entry, same die), peakClockMHz (1455 boost), and memoryBandwidthGBs (653, from HBM2 at 850MHz effective on a 3,072-bit bus) are corroborated consistently across independent GPU databases, not individually restated in NVIDIA's press release. fp32TFLOPS (14.9) and fp64TFLOPS (7.4) are likewise secondary-sourced and not from an NVIDIA table; they're internally consistent with V100 SXM2's own official 15.7/7.8 TFLOPS scaled down by Titan V's slightly lower boost clock (1455MHz vs V100's 1530MHz) — flagged here as the least-certain figures on this entry, worth a final check against a primary source if precision matters. fp16TFLOPS (110) is the one Tensor figure NVIDIA published directly. tdpWatts (250) is well-corroborated but not NVIDIA-stated.\n2026-09-15 — added computeUnitCount (80 SMs), consistent with the shaderCoreCount (5,120) and matrixCoreCount (640) already recorded here: a GV100 SM carries 64 FP32 cores and 8 Tensor Cores, so both figures independently give 80 SMs. That is the same 80-of-84-SM GV100 configuration NVIDIA's Volta whitepaper documents for Tesla V100 (see this collection's V100 entry) — TITAN V is the same die and the same SM count, with consumer-side memory (12GB HBM2) and clocks.\n"
    },
    {
      "id": "nvidia-titan-x-pascal",
      "vendor": "nvidia",
      "name": "TITAN X (Pascal)",
      "category": "consumer",
      "architecture": "Pascal",
      "releaseYear": 2016,
      "launchDate": "Aug 2, 2016",
      "targetId": "6.1",
      "vramGB": 12,
      "memoryType": "GDDR5X",
      "memoryBandwidthGBs": 480,
      "fp32TFLOPS": 11,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 28,
      "shaderCoreCount": 3584,
      "peakClockMHz": 1531,
      "tdpWatts": 250,
      "pcieGen": "PCIe Gen3 x16",
      "msrpUSD": 1200,
      "notes": "Added 2026-09 alongside GTX 1080 Ti, TITAN V, and TITAN RTX — NVIDIA's first Pascal-generation TITAN, explicitly marketed at launch for \"deep learning and artificial intelligence\" (per NVIDIA's own GTX 1080 Ti press release, which describes this card as \"designed for deep learning and artificial intelligence\"). Not to be confused with the earlier, differently-architected Maxwell \"GeForce GTX TITAN X\" (2015) — this is the GP102-based \"(Pascal)\" refresh, hence the parenthetical in both NVIDIA's own naming and this entry's name field.\nCUDA core count (3,584), memory config (12GB GDDR5X at 10Gbps, 384-bit), memory bandwidth (480 GB/sec), and FP32 performance (\"11 TFLOPS of power\") are quoted directly from NVIDIA's own official product page (nvidia.com/en-us/geforce/news/gfecnt/nvidia-titan-x-pascal-available- august-2nd/), read 2026-09. peakClockMHz (1531, boost) is from the same page's \"boosts to 1.5GHz out of the box\" combined with the widely-corroborated exact figure across independent GPU databases; tdpWatts (250) and msrpUSD (1200, the original announced price) are not stated on that specific NVIDIA page but are consistently corroborated across contemporaneous press coverage of the same Jul 21 2016 announcement. No Tensor/matrix-throughput fields — Pascal predates Tensor Cores (introduced with Volta the following year).\n2026-09-15 — added computeUnitCount (28 SMs): the same GP102 configuration as this collection's GTX 1080 Ti entry — 3,584 CUDA cores / 128 FP32 cores per Pascal consumer SM = 28 of the die's 30 SMs enabled. NVIDIA publishes the core total, not the SM count.\n"
    },
    {
      "id": "nvidia-v100-sxm2-32gb",
      "vendor": "nvidia",
      "name": "V100 SXM2 32GB",
      "category": "datacenter",
      "architecture": "Volta",
      "releaseYear": 2017,
      "launchDate": "May 2017 (original 16GB V100); 32GB SXM2 variant added 2018",
      "targetId": "7.0",
      "vramGB": 32,
      "memoryType": "HBM2",
      "memoryBandwidthGBs": 900,
      "eccSupport": true,
      "fp64TFLOPS": 7.8,
      "fp32TFLOPS": 15.7,
      "fp16TFLOPS": 125,
      "computeUnitLabel": "Streaming Multiprocessors",
      "computeUnitCount": 80,
      "shaderCoreCount": 5120,
      "matrixCoreCount": 640,
      "tdpWatts": 300,
      "formFactor": "SXM2",
      "interconnect": "NVLink",
      "interconnectBandwidthGBs": 300,
      "notes": "Added 2026-09 to fill a gap this site's own README had flagged (oldest NVIDIA entry was 2020's A100; V100 is still a common ML training/inference GPU in the field). All figures sourced directly from NVIDIA's official V100 datasheet (images.nvidia.com/content/technologies/volta/pdf/ volta-v100-datasheet-update-us-1165301-r5.pdf, \"V100 SXM2\" column, footer dated Jan20), read via pdftotext after WebFetch could not parse this PDF's text layer directly (same issue noted on this site's A100/B300 entries).\nfp16TFLOPS (125) is the datasheet's single \"Tensor Performance\" row — Volta's 1st-gen Tensor Cores predate structured sparsity (introduced with Ampere), so there is no dense/sparse split and no fp16TFLOPSSparse value; this is the only Tensor-Core throughput row Volta publishes (no separate BF16/TF32/FP8/INT8 Tensor figures — BF16 and TF32 as Tensor-Core formats and FP8 didn't exist until later architectures). fp16TFLOPSVector (packed FP16 via CUDA cores) isn't in this datasheet either, so it's omitted rather than guessed at 2x fp32TFLOPS.\nThis is the SXM2 form factor at 300W; the same datasheet lists a PCIe variant (V100 PCIe, 250W max power, 7/14/112 TFLOPS FP64/FP32/Tensor, 16 or 32GB, 900GB/s bandwidth, 32GB/sec PCIe Gen3 interconnect instead of NVLink) and a V100S PCIe refresh (8.2/16.4/130 TFLOPS, 32GB HBM2 only, 1134GB/s bandwidth, still 250W) — not modeled as separate entries here, noted for completeness. interconnectBandwidthGBs (300) is NVLink specifically; the SXM2 module also has a PCIe Gen3 host link not given a separate bandwidth figure in this table. No transistor count, process node, or MSRP in this datasheet — V100 shipped through OEM/server partners, not direct retail, and this table doesn't list die-level figures the way NVIDIA's longer Volta architecture whitepaper might.\n2026-09-15 — added computeUnitCount (80 SMs) from NVIDIA's own Volta architecture whitepaper (WP-08608-001_v1.1), which states that \"the Tesla V100 accelerator uses 80 SMs\" of the full GV100's 84, and whose Table 1 lists the matching 5,120 FP32 cores and 640 Tensor Cores already recorded here (64 FP32 cores and 8 Tensor Cores per SM). Read via pdftotext after WebFetch failed on the PDF's binary structure — the same workaround this collection's RTX 4090 entry documents.\n"
    }
  ]
}