NVIDIA

GB200

NVL72

The NVIDIA GB200 NVL72 is a high-performance GPU variant designed for data-intensive workloads in the datacenter. It targets enterprise and research markets, offering exceptional computational power for AI and machine learning tasks. As part of the Ampere architecture, it features advanced tensor cores and high memory bandwidth, making it suitable for large-scale model training and inference.

GB200 NVL72 — illustration of the card's form factor
Memory
HBM3E
Architecture
Blackwell
Form Factor
rack-scale

Provider Marketplace

Cheapest
$10.50/hour
Starting from
Best Value
Awaiting listings
No additional provider yet
Enterprise Choice
Awaiting listings
No additional provider yet

All Cloud Providers

1 Options available
CoreWeave logo
CoreWeaveCheapest
On-DemandNORTH AMERICA
$10.50/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Architecture

MicroarchitectureBlackwell
Matrix EngineTransformer Engine (2nd generation)
Transformer EngineYes (2nd generation)
Sparse AccelerationSupported (sparse/dense; dense is one-half sparse spec shown)
Dynamic PrecisionSupported (FP4, FP8, FP6, FP16, BF16, TF32, FP32, FP64, INT8)

Memory & VRAM

Memory TypeHBM3E
Compressiondedicated decompression engines
Memory Pooling72-GPU NVLink domain that acts as a single, massive GPU

Connectivity & Scaling

InterconnectNVLink
Generationfifth-generation NVIDIA NVLink
TopologyNVLink Switch System
Scale-OutNVIDIA Quantum-X800 InfiniBand; NVIDIA Spectrum-X800 Ethernet

Power & Efficiency

Thermal Limitsrack-scale, liquid-cooled design
Efficiency25x more performance at the same power (vs H100)

Physical Design

Form Factorrack-scale
Coolingliquid-cooled
Rack DensityIncreases compute density

Thermals & Cooling

Liquid Coolingrack-scale, liquid-cooled design
DC Heatreduces a data center’s carbon footprint and energy consumption

Server & Deployment

Rack-Scalerack-scale, liquid-cooled design

System Compatibility

CPU PairingTwo NVIDIA Blackwell GPUs connected to one NVIDIA Grace CPU via NVLink-C2C (per GB200 Grace Blackwell Superchip)
OS CompatNVIDIA Mission Control NVIDIA AI Enterprise NVIDIA DGX OS / Ubuntu

Benchmarks & Throughput

Transformer Throughput

GB200 NVL72 introduces a second-generation Transformer Engine which enables FP4 AI.

Training Benchmarks

GB200 NVL72 offers a faster second-generation Transformer Engine with FP8 precision and delivers 4x faster training for large language models versus H100.

Inference Benchmarks

GB200 NVL72 delivers 30x faster real-time LLM inference performance for trillion-parameter language models versus NVIDIA H100.

Scaling Efficiency

72-GPU NVLink domain acting as a single massive GPU with NVLink Switch System providing 130 TB/s of low-latency GPU communications.

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksGB200 NVL72 uses NVLink and liquid cooling to overcome communication bottlenecks
ParallelismScale-up NVLink domain (72-GPU acts as single massive GPU) and scale-out via InfiniBand/Ethernet/DPU networking

Workload Readiness

HPC / Simulation

Designed for AI and high-performance computing (HPC) workloads

Real-Time Serving

Includes a second-generation Transformer Engine enabling FP4 AI for real-time LLM inference

Market Authority

Supercomputer Usage

exascale computer in a single rack

Key Strengths

The GB200 NVL72 excels at handling large-scale AI and machine learning tasks, offering superior performance in model training and inference. Its advanced architecture and high memory bandwidth make it stand out for demanding computational workloads.

Limitations

Potential limitations include high power consumption and cooling requirements. Availability may be constrained by demand and production capacity. Users should ensure compatibility with existing infrastructure and consider the cost implications of deploying such high-performance hardware.

Expert Insight

The GB200 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.