Accelerator
Instinct MI455X
AMD CDNA 5 gpu (per accelerator) summary for training, inference, and roofline-style performance analysis.
Notes
- All compute and memory figures are per GPU vendor peaks, not sustained measurements or 72-GPU Helios rack totals. Launch is AMD's listed product launch date.
- Compute values use the July 2026 datasheet's TFLOPS/TOPS table; AMD's product webpage rounds these figures to PFLOPS/POPS.
- BLOCK FP4, BLOCK FP6, and BLOCK FP8 represent OCP MXFP4, MXFP6, and MXFP8 respectively. No separate plain FP4, plain FP6, INT4, or INT32 peaks are published in the cited datasheet.
- Only BF16, FP16, and INT8 have explicit structured-sparsity peaks in the datasheet. Missing sparse entries mean unreported, not necessarily unsupported.
- BF16 and FP16 entries are matrix peaks. FP16 vector peak is 315 TFLOPS; FP32 matrix/vector peaks are both 315 TFLOPS and FP64 matrix/vector peaks are both 5 TFLOPS.
- The GPU has 12 HBM4 stacks and 192 MB L2 cache. AMD lists 256 Work Group Processors, eight accelerated compute dies, two I/O dies, and a 2.4 GHz peak engine clock.
- The datasheet lists a 256 GB/s CPU-to-GPU interconnect. Helios combines four GPUs per compute tray and 72 per rack; these aggregate capacities are not used in the table or roofline plot.
- Direct liquid cooling is specified, but neither the cited product page nor datasheet provides a per-GPU TDP or TBP.
Sources
- AMD MI455X GPU datasheet (July 2026) Accessed 2026-09-05
- AMD MI455X product specifications Accessed 2026-09-05
- AMD CDNA 5 and Helios architecture overview Accessed 2026-09-05