AMD · 2020-11-16

Instinct MI100

The AMD Instinct MI100 accelerator is designed to power HPC workloads and speed up time-to-discovery. It is built on the AMD CDNA architecture.

Instinct MI100 — illustration of the card's form factor
VRAM
32 GB
FP32 TFLOPS
23.1 TFLOPS
CUDA Cores
7,680
TDP
300 W

Compute Performance

FP6411.5 TFLOPS
FP3223.1 TFLOPS
FP16184.6 TFLOPS
BF1692.3 TFLOPS
INT492.3 TOPS

Architecture

MicroarchitectureCDNA
Process NodeTSMC 7nm FinFET
Compute Units120 CUs
CUDA Cores7680
Boost Clock1502 MHz

Memory & VRAM

Memory TypeHBM2
Total Capacity32 GB
Bandwidth1.2 TB/s
Bus Width4096-bit
ECC SupportYes (Full-Chip)

Connectivity & Scaling

InterconnectInfinity Fabric
IB Bandwidth92 GB/s
PCIe InterfacePCIe 4.0, PCIe 3.0 xx16

Power & Efficiency

TDP300 W
PSU RequiredPCIe Powered
ConnectorsPCIe Powered
Thermal LimitsPassive

Physical Design

Form FactorPCIe
Slot WidthDouble Slot
CoolingPassive

System Compatibility

Required PCIePCIe® 4.0 x16, PCIe® 3.0 x16
MotherboardPCIe® Add-in Card (PCIe® 4.0 x16, PCIe® 3.0 x16)
Rack PowerThermal Design Power (TDP) 300W Peak; External Power Connectors: PCIe Powered
OS CompatRHEL x86 64-Bit (Radeon™ Software for Linux® version 25.35 for RHEL 10.1; Radeon™ Software for Linux® version 25.35 for RHEL 9.7), Ubuntu x86 64-Bit (Radeon™ Software for Linux® version 25.35 for Ubuntu 24.04.4 HWE; Radeon™ Software for Linux® version 25.35 for Ubuntu 22.04.5 HWE), SLED/SLES 15 (Radeon™ Software for Linux® version 25.35 for SLED/SLES 15 SP7)

Benchmarks & Throughput

Scaling Efficiency

AMD Infinity Fabric links; up to 4-GPU hive topology; Peak Infinity Fabric Link Bandwidth 92 GB/s; Infinity Fabric Links 3

Multi-GPU Scalability

Scaling Characteristics

ParallelismSupported Technologies AMD Infinity Architecture , AMD CDNA™ Architecture , AMD ROCm™ - Ecosystem without Borders

Workload Readiness

HPC / Simulation

AMD Instinct™ MI100 accelerators are designed to power HPC workloads and speed time-to-discovery.

Scientific Computing

11.5 TFLOPs

Market Authority

Key Strengths

The MI100 excels in AI and HPC workloads with its high FP64 performance.

  • ·FP64 Performance: Offers strong double-precision performance for scientific computing.
  • ·AI Training: Optimized for AI training with high throughput.
  • ·PCIe 4.0: Leverages PCIe 4.0 for faster data transfer rates.

Limitations

The MI100 has some limitations in terms of availability and specific workload optimizations.

  • ·Availability: May have limited availability compared to NVIDIA counterparts.
  • ·Software Ecosystem: Less mature software ecosystem compared to NVIDIA CUDA.

Expert Insight

The Instinct MI100 represents a powerful alternative for diversified workloads. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.