AMD · 2023-12-06

Instinct MI210

PCIe Gen4 Passive Accelerator

The AMD Instinct MI210 PCIe Gen4 Passive Accelerator is a compute workhorse optimized for accelerating single precision and double-precision HPC-class systems. It offers Exascale-Class Technologies, Purpose-built Accelerators for HPC & AI Workloads, and Innovations Delivering Performance Leadership.

Instinct MI210 PCIe Gen4 Passive Accelerator — illustration of the card's form factor
VRAM
64GB
FP32 TFLOPS
22.6 TFLOPS
CUDA Cores
6,656
TDP
300 W

Provider Marketplace

Cheapest
$0.16/hour
Starting from
Best Value
Awaiting listings
No additional provider yet
Enterprise Choice
Awaiting listings
No additional provider yet

All Cloud Providers

1 Options available
RedSwitches logo
RedSwitchesCheapest
On-DemandLondonUnited Kingdom
$0.16/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP6422.6 TFLOPS
FP3222.6 TFLOPS
FP16181 TFLOPS
BF16181 TFLOPS
INT8181 TOPS
INT4181 TOPS

Architecture

MicroarchitectureCDNA2
Process Node6nm FinFet
Compute Units104 CUs
CUDA Cores6656
Tensor CoresMatrix Cores 416
Matrix EngineAMD Matrix Core Technology
Boost Clock1700 MHz
Dynamic PrecisionSupported (BF16/FP16/FP32/INT8/INT4)

Memory & VRAM

Memory TypeHBM2e
Total Capacity64GB
Bandwidth1.6 TB/s
Bus Width4096-bit
ECC SupportYes (Full-Chip)

Connectivity & Scaling

InterconnectInfinity Fabric
Generation3rd Gen AMD Infinity Architecture
IB Bandwidth100 GB/s
PCIe InterfacePCIe® 4.0 xx16
TopologyDual | Quad Hives
P2P MemoryYes (Peer-to-Peer / Coherency Enabled)

Virtualization

SR-IOVYes (Passthrough Only)
GPU SharingSR-IOV (Passthrough Only)

Power & Efficiency

TDP300 W
Peak Power300
PSU RequiredEPS12V, 8-pin
ConnectorsEPS12V, 8-pin
Thermal LimitsPassively Cooled

Physical Design

Form FactorPCIe
FHFLYes
Slot WidthDual-slot
Dimensions4.5” x 10.5” (11.43 CM x 26.67 CM)
CoolingPassive

Software Ecosystem

ROCmROCm 6
PyTorchPyTorch
TensorFlowTensorFlow
Compiler StackOpenMP | HIP | OpenCL™ | Python

System Compatibility

Required PCIePCIe Gen4
MotherboardPCIe® x16 add-in card, Full-Height Full-Length dual-slot
Rack Power300W TDP (EPS12V, 8-pin)
OS CompatLinux 64 Bit (AMD ROCm Compatible)

Benchmarks & Throughput

Scaling Efficiency

AMD Instinct™ MI210 CDNA 2 technology-based accelerators include three Infinity Fabric™ links providing up to 300 GB/s peak theoretical GPU to GPU or Peer-to-Peer (P2P) bandwidth performance per GPU card.

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksPCIe Gen4 CPU->GPU up to 64 GB/s per card; Infinity Fabric P2P up to 300 GB/s per GPU card; aggregate GPU card I/O peak bandwidth up to 364 GB/s.
ParallelismCoherency enabled (Dual | Quad Hives); AMD ROCm compatible; SR-IOV support (Passthrough Only); supports OpenMP | HIP | OpenCL programming models.

Workload Readiness

LLM Training

true

Vision Training

true

Reinforcement Learning

true

HPC / Simulation

true

Scientific Computing

true

Market Authority

Research Citations

We thank the Computational Infrastructure for Geodynamics (http://geodynamics.org) which is funded by the National Science Foundation under awards EAREcosystem without Borders

Key Strengths

Excels in high-performance and AI workloads.

  • ·AI Training: Optimized for large-scale AI model training.
  • ·HPC Performance: Delivers strong performance in scientific computing tasks.
  • ·Data Analytics: Efficient for large-scale data processing and analytics.

Limitations

Some limitations in software ecosystem compared to NVIDIA.

  • ·Software Ecosystem: Less mature software stack compared to NVIDIA CUDA.
  • ·Availability: May have limited availability in certain regions.

Expert Insight

The Instinct MI210 represents a powerful alternative for diversified workloads. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.