AMD · 2025-01-01

Instinct MI300A

APU

The AMD Instinct MI300A APU is a breakthrough discrete accelerated processing unit designed for high-performance computing and AI applications. It integrates 24 AMD 'Zen 4' x86 CPU cores with 228 AMD CDNA™ 3 high-throughput GPU compute units and 128 GB of unified HBM3 memory.

Instinct MI300A APU — illustration of the card's form factor
VRAM
128 GB
FP32 TFLOPS
122.6 TFLOPS
CUDA Cores
14,592
TDP
550 W

Compute Performance

FP6461.3 TFLOPS
FP32122.6 TFLOPS
TF32490.3 TFLOPS
FP16980.6 TFLOPS
BF16980.6 TFLOPS
FP81961.2 TFLOPS
INT81961.2 TOPS

Architecture

MicroarchitectureCDNA3
Process NodeTSMC 5nm | 6nm FinFET
Transistors146 Billion
Compute Units228 CUs
CUDA Cores14592
Tensor Cores912 Matrix Cores
Boost Clock2100 MHz
Sparse AccelerationSupported (structured sparsity)
Dynamic PrecisionSupported (FP64/FP32/TF32/FP16/BF16/FP8/INT8)

Memory & VRAM

Memory TypeHBM3
Total Capacity128 GB
Bandwidth5.3 TB/s
Bus Width8192-bit
ECC SupportYes (Full-Chip)
Unified MemoryYes (single shared address space between CPU and GPU)

Connectivity & Scaling

InterconnectInfinity Fabric
Generation4th Gen AMD Infinity architecture
IB Bandwidth1 TB/s
PCIe InterfacePCIe 5.0 xx16
TopologyInfinity Fabric peer-to-peer (typical 4-APU: six interfaces dedicated)
Max GPUs/Node4
Scale-Out400 Gbps Ethernet or InfiniBand
P2P MemoryUnified shared address space between CPU and GPU

Virtualization

SR-IOVSupported
GPU SharingSR-IOV (up to 3 partitions)
Virt EfficiencyNear bare-metal (vendor claim)

Power & Efficiency

TDP550 W
Peak Power760
Thermal LimitsPassive & Active cooling

Physical Design

Form FactorAPU SH5 socket
CoolingPassive & Active

Thermals & Cooling

Liquid Coolingtrue
DC Heat550W (air & liquid cooling); 760W (liquid cooling)

Software Ecosystem

ROCmAMD ROCm 6
PyTorchSupported
TensorFlowSupported
JAXSupported
Triton ServerSupported
DockerSupported (GPU software containers available)
Compiler StackSupported (mature drivers and compilers)
Driver StabilityMature drivers

Server & Deployment

OEM Availabilityavailable to enterprise data centers through platforms offered by our solution partners
Preconfiguredavailable to enterprise data centers through platforms offered by our solution partners
Ref ArchitecturesExample server architecture with four interconnected APUs

System Compatibility

CPU PairingIntegrated 24 AMD ‘Zen 4’ x86 CPU cores
NUMA128 GB of unified HBM3 memory that presents a single shared address space to CPU and GPU
Required PCIePCIe® 5.0
MotherboardAPU SH5 socket
Rack Power550W (air & liquid cooling); 760W (liquid cooling)

Benchmarks & Throughput

Structured Sparsity

Supported (document lists peak performance 'with Structured Sparsity' for multiple datatypes)

Scaling Efficiency

Typical 4-APU configuration: six interfaces dedicated to inter-GPU Infinity Fabric connectivity for a total of 384 GB/s peer-to-peer connectivity per APU

Multi-GPU Scalability

Scaling Characteristics

Network Bottlenecksperformance bottlenecks from the narrow interfaces between CPU and GPU
ParallelismSR-IOV (up to 3 partitions); ROCm software support for parallelizing across multiple GPUs and servers

Workload Readiness

LLM Training

Supported

LLM Inference

Supported

HPC / Simulation

Supported

Scientific Computing

Supported

Real-Time Serving

Supported

Market Authority

Supercomputer Usage

El Capitan system at Lawrence Livermore National Labs (expected to become the next world’s fastest supercomputer)

Key Strengths

Excels in mixed workloads requiring both CPU and GPU resources.

  • ·AI Workloads: Optimized for AI training and inference tasks.
  • ·HPC Applications: Strong performance in high-performance computing scenarios.
  • ·Energy Efficiency: Combines CPU and GPU for improved energy efficiency.

Limitations

Limited by platform-specific requirements and availability.

  • ·Platform Specific: Requires compatible server infrastructure for deployment.
  • ·Availability: May have limited availability in certain regions or markets.

Expert Insight

The Instinct MI300A represents a powerful alternative for diversified workloads. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.