AMD · 2025-01-01

Instinct MI300X

The AMD Instinct MI300X discrete GPU is based on next-generation AMD CDNA 3 architecture, featuring 304 high-throughput compute units, AI-specific functions, and 192 GB of HBM3 memory. It offers outstanding performance for demanding AI and HPC applications, with a focus on generative AI, machine learning, and inferencing.

Instinct MI300X — illustration of the card's form factor
VRAM
192 GB
FP32 TFLOPS
163.4 TFLOPS
CUDA Cores
19,456
TDP
750 W

Provider Marketplace

Cheapest
$1.71/hour
Starting from
Best Value
$2.99/hour
Starting from
Enterprise Choice
$3.45/hour
Starting from

All Cloud Providers

4 Options available
TensorWave logo
TensorWaveCheapest
On-Demand
$1.71/ hour
Estimated Cost
Provision
DigitalOcean logo
Reserved · commitment
$1.91/ hour
Estimated Cost
Provision
Hot Aisle logo
On-Demand
$2.99/ hour
Estimated Cost
Provision
Crusoe Cloud logo
On-Demand
$3.45/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP6481.7 TFLOPS
FP32163.4 TFLOPS
TF32653.7 TFLOPS
FP161307.4 TFLOPS
BF161307.4 TFLOPS
FP82614.9 TFLOPS
INT82614.9 TOPS

Architecture

MicroarchitectureCDNA3
Process NodeTSMC 5nm | 6nm FinFET
Transistors153 Billion
Compute Units304 CUs
CUDA Cores19456
Tensor CoresMatrix Cores, 1216
Matrix EngineAMD Matrix Core
Boost Clock2100 MHz
Sparse AccelerationSupported (structured sparsity)
Dynamic PrecisionSupported (FP8/FP16/TF32/FP32/FP64/BF16/INT8)

Memory & VRAM

Memory TypeHBM3
Total Capacity192 GB
Bandwidth5.3 TB/s
Bus Width8192-bit
ECC SupportYes (Full-Chip)
Unified MemoryYes (shared coherently between CPUs and GPUs)

Connectivity & Scaling

InterconnectInfinity Fabric
Generation4th Gen Infinity Architecture
IB Bandwidth128 GB/s
PCIe InterfaceGen 5 xx16
TopologyRing
Max GPUs/Node8
P2P MemoryShared coherently between CPUs and GPUs

Virtualization

SR-IOVSupported
GPU SharingSR-IOV; Coherent shared memory and caches
Virt EfficiencyNear bare-metal (vendor claim)

Power & Efficiency

TDP750 W
Peak Power750
Connectors54V UBB
Thermal LimitsCooling Passive OAM

Physical Design

Form FactorOAM Module
CoolingPassive OAM

Thermals & Cooling

DC Heat750W

Software Ecosystem

ROCmROCm 6
PyTorchSupported
TensorFlowSupported
JAXSupported
Triton ServerSupported
Compiler StackROCm compilers
Driver StabilityMature drivers

Server & Deployment

PreconfiguredThe discrete MI300X is sold as an AMD Instinct Platform with eight accelerators interconnected on an AMD Universal Base Board (UBB 2.0) with industry-standard HGX host connectors.
DGX/HGXThe discrete MI300X is sold as an AMD Instinct Platform with eight accelerators interconnected on an AMD Universal Base Board (UBB 2.0) with industry-standard HGX host connectors.

System Compatibility

Required PCIePCIe® 5.0 x16

Benchmarks & Throughput

Structured Sparsity

Supported

Scaling Efficiency

Seven Infinity Fabric links for full connectivity between eight GPUs in a ring (128 GB/s bidirectional per link)

Multi-GPU Scalability

Scaling Characteristics

ParallelismCoherent shared memory and caches; SR-IOV (up to 8 partitions); Infinity Fabric links for GPU-to-GPU connectivity (7 links in a ring)

Workload Readiness

LLM Training

Supported

LLM Inference

Supported

HPC / Simulation

Supported

Scientific Computing

Supported

Market Authority

Supercomputer Usage

Document states MI300X is already powering the fastest exaFLOP-class HPC and that AMD Instinct GPUs have been deployed on Green500 supercomputers.

Research Citations

Document includes Top500 and Green500 references.

Key Strengths

The MI300X excels in AI and HPC workloads with its advanced architecture.

  • ·AI Training: Optimized for large-scale AI model training with high throughput.
  • ·HPC Performance: Delivers exceptional performance for high-performance computing tasks.
  • ·Memory Bandwidth: Features high memory bandwidth for data-intensive applications.

Limitations

The MI300X has some limitations in terms of availability and specific workload optimizations.

  • ·Availability Constraints: May have limited availability due to high demand and production constraints.
  • ·Workload Optimization: While strong in AI, may not be as optimized for certain niche workloads compared to competitors.

Expert Insight

The Instinct MI300X represents a powerful alternative for diversified workloads. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.