M40 24GB
NVIDIAdatacenter| Overview | |
|---|---|
| Architecture | Maxwell (2015) |
| Launch date | Nov 2015 (12GB, datasheet dated Jan16); 24GB variant datasheet dated Mar16 |
| Compute capability | 5.2 |
| Memory | |
| VRAM | 24 GB GDDR5 |
| Memory bandwidth | 288 GB/s |
| Compute — vector | |
| FP64 | 0.2 TFLOPS |
| FP32 | 7 TFLOPS |
| Cores & clocks | |
| Streaming Multiprocessors | 24 |
| Shader cores | 3,072 |
| Board & system | |
| TDP | 250 W |
| Form factor | PCIe Full Height, Dual Slot |
| PCIe | PCIe Gen3 |
Compatible CUDA Toolkit versions
Added 2026-09 alongside K80 and older AMD Radeon Instinct cards to extend NVIDIA coverage further back — M40 was a widely-deployed early deep learning training GPU (Caffe/Torch era) and predates this site's previous oldest NVIDIA entry (2017's V100) by two years. All figures sourced directly from NVIDIA's official Tesla M40 datasheet, 24GB variant (images.nvidia.com/content/tesla/pdf/78071_Tesla_M40_24GB_Print_Datasheet_LR.PDF, footer dated MAR16), read via pdftotext. NVIDIA also published an otherwise-identical 12GB datasheet (nvidia-teslam40-datasheet.pdf, dated Jan16) with the same architecture/CUDA-core-count/TFLOPS/bandwidth/power figures — this entry represents the 24GB SKU, matching this site's general preference for the higher-memory variant where both exist (see V100 SXM2 32GB's notes for the same reasoning).
fp32TFLOPS (7) is explicitly "with NVIDIA GPU Boost" per the datasheet's own Features list. eccSupport is intentionally omitted (not false) — unlike this site's other Tesla-era entries (K80, P100, V100), M40's official spec table has no ECC row at all, and Maxwell GM200's GDDR5 (rather than HBM2) memory controller doesn't get the same blanket ECC treatment those datasheets call out, so leaving it unspecified is more accurate than guessing either way. No Tensor/matrix-throughput fields — Maxwell predates Tensor Cores entirely (introduced with Volta), and this datasheet gives only a single FP32 rate, not a separate FP16 vector figure. No transistor count, process node, or MSRP — M40 shipped through OEM/server partners, not direct retail.
2026-09-15 — added computeUnitCount (24 SMs) from Table 1 of NVIDIA's Volta architecture whitepaper (WP-08608-001_v1.1), which lists GM200 (Maxwell) at 24 SMs, 128 FP32 cores per SM and 3,072 FP32 cores per GPU — the same 3,072 already in shaderCoreCount here. Maxwell's own name for the block is SMM (as Kepler's is SMX); computeUnitLabel says "Streaming Multiprocessors" for consistency with the rest of this collection's NVIDIA entries. matrixCoreCount stays empty — no Tensor Cores on Maxwell.