Accelerator

Instinct MI455X

AMD CDNA 5 gpu (per accelerator) summary for training, inference, and roofline-style performance analysis.

Back to Accelerator Catalog

Vendor
AMD
Architecture
CDNA 5
Unit
GPU (per accelerator)
Form factor
Enhanced Accelerator Module (EAM)
Launch
2026-07-23
Memory
432 GB HBM4
HBM bandwidth
23.3 TB/s
BF16 peak
5.03 PFLOPS
BF16 sparse peak
10.07 PFLOPS
FP16 peak
5.03 PFLOPS
FP16 sparse peak
10.07 PFLOPS
FP8 dense peak
20.13 PFLOPS
FP8 sparse peak
n/a
FP4 dense peak
n/a
FP4 sparse peak
n/a
BLOCK FP4 dense peak
40.27 PFLOPS
BLOCK FP6 dense peak
20.13 PFLOPS
BLOCK FP8 dense peak
20.13 PFLOPS
FP32 peak
315 TFLOPS
FP64 peak
5 TFLOPS
INT8 peak
5.03 POPS
INT8 sparse peak
10.07 POPS
Interconnect
UALoE (UALink over Ethernet) - 3.6 TB/s peak bidirectional scale-up per GPU; 600 GB/s peak bidirectional scale-out per GPU
Power
Not published in cited AMD specifications
Software stack
ROCm

Notes

  • All compute and memory figures are per GPU vendor peaks, not sustained measurements or 72-GPU Helios rack totals. Launch is AMD's listed product launch date.
  • Compute values use the July 2026 datasheet's TFLOPS/TOPS table; AMD's product webpage rounds these figures to PFLOPS/POPS.
  • BLOCK FP4, BLOCK FP6, and BLOCK FP8 represent OCP MXFP4, MXFP6, and MXFP8 respectively. No separate plain FP4, plain FP6, INT4, or INT32 peaks are published in the cited datasheet.
  • Only BF16, FP16, and INT8 have explicit structured-sparsity peaks in the datasheet. Missing sparse entries mean unreported, not necessarily unsupported.
  • BF16 and FP16 entries are matrix peaks. FP16 vector peak is 315 TFLOPS; FP32 matrix/vector peaks are both 315 TFLOPS and FP64 matrix/vector peaks are both 5 TFLOPS.
  • The GPU has 12 HBM4 stacks and 192 MB L2 cache. AMD lists 256 Work Group Processors, eight accelerated compute dies, two I/O dies, and a 2.4 GHz peak engine clock.
  • The datasheet lists a 256 GB/s CPU-to-GPU interconnect. Helios combines four GPUs per compute tray and 72 per rack; these aggregate capacities are not used in the table or roofline plot.
  • Direct liquid cooling is specified, but neither the cited product page nor datasheet provides a per-GPU TDP or TBP.

Sources