AMD · 2021-09-21

Instinct MI200

The AMD Instinct MI200 is a high-performance GPU accelerator based on the 2nd Gen AMD CDNA architecture. It offers industry-leading double precision performance for HPC workloads, with up to 47.9 TFLOPS peak FP64 performance. The MI200 is optimized for AI and machine learning workloads, supporting a full range of mixed precision operations.

Instinct MI200 — illustration of the card's form factor
VRAM
128GB
CUDA Cores
14,080
TDP
500 W
Memory
HBM2e

Compute Performance

FP6447.9 TFLOPS
FP16383 TFLOPS
BF16383 TFLOPS

Architecture

MicroarchitectureCDNA 2
Process Node6nm FinFet
Matrix EngineMatrix Core technology
Dynamic PrecisionSupported (INT4/INT8/FP8/FP16/BF16/FP32/FP64)

Memory & VRAM

Memory TypeHBM2e
Total Capacity128GB
Bandwidth3.2 TB/s
Bus Width8,192 bits
ECC SupportYes (Full-chip)

Physical Design

Form FactorOAM
CoolingPassive & Liquid

Software Ecosystem

ROCmROCm 5.0
PyTorchPyTorch
TensorFlowTensorFlow
Compiler StackOpenMP | HIP | OpenCL | Python

Server & Deployment

OEM AvailabilityFind a partner offering AMD Instinct accelerator-based solutions.
Rack-Scaleat any scale—from single-server solutions up to the world’s largest, Exascale-class supercomputers.1

System Compatibility

CPU PairingOptimized 3rd Gen AMD EPYC™ CPU
Required PCIePCIe® Gen 4
MotherboardOAM
OS CompatLinux 64 Bit

Benchmarks & Throughput

Training Benchmarks

the MI200 accelerator is the first data center GPU to deliver 383 teraflops of theoretical mixed precision FP16 performance for deep learning training.

Scaling Efficiency

AMD Instinct MI200 series OAM accelerators with advanced peer-to-peer I/O connectivity through a maximum of eight AMD Infinity Fabric™ links deliver up to 800 GB/s I/O bandwidth performance.

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksInfinity Fabric peer-to-peer I/O up to 800 GB/s per OAM card; HBM2e memory bandwidth up to 3.2 TB/s
ParallelismOpenMP | HIP | OpenCL™ | Python

Workload Readiness

LLM Training

Supported (peak FP16 for deep learning training: 383 teraflops)

Vision Training

Supported

Reinforcement Learning

Supported

HPC / Simulation

Supported

Scientific Computing

Supported

Market Authority

Supercomputer Usage

Selected for the first U.S. Exascale supercomputer / powers some of the world’s top supercomputers

Community Benchmarks

HPCG 3.0, HPL, HPL-AI, PyFR, OpenFOAM, AMBER, GROMACS, LAMMPS, NAMD (benchmarks listed on product page)

Enterprise Cases

["Biznet Gio Scales Cloud Performance for AI Era with AMD","LiquidMetal AI Delivers Cost-Effective GenAI Performance with AMD","310 AI Advances Protein Design with AMD","MindWalk™ Accelerates AI Drug Discovery with AMD"]

Key Strengths

The MI200 excels in high-performance computing and AI training tasks.

  • ·HPC Performance: Optimized for high-performance computing with advanced matrix operations.
  • ·AI Training: Efficient for large-scale AI model training with high throughput.
  • ·Energy Efficiency: Designed for improved performance per watt with MCM architecture.

Limitations

The MI200 series has some limitations in terms of availability and compatibility.

  • ·Availability: Limited availability in certain regions and platforms.
  • ·Compatibility: Requires specific infrastructure for optimal deployment.

Expert Insight

The Instinct MI200 represents a powerful alternative for diversified workloads. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.