NVIDIA
GB200
NVL72
The NVIDIA GB200 NVL72 is a high-performance GPU variant designed for data-intensive workloads in the datacenter. It targets enterprise and research markets, offering exceptional computational power for AI and machine learning tasks. As part of the Ampere architecture, it features advanced tensor cores and high memory bandwidth, making it suitable for large-scale model training and inference.

Provider Marketplace
All Cloud Providers
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Architecture
Memory & VRAM
Connectivity & Scaling
Power & Efficiency
Physical Design
Thermals & Cooling
Server & Deployment
System Compatibility
Benchmarks & Throughput
Transformer Throughput
GB200 NVL72 introduces a second-generation Transformer Engine which enables FP4 AI.
Training Benchmarks
GB200 NVL72 offers a faster second-generation Transformer Engine with FP8 precision and delivers 4x faster training for large language models versus H100.
Inference Benchmarks
GB200 NVL72 delivers 30x faster real-time LLM inference performance for trillion-parameter language models versus NVIDIA H100.
Scaling Efficiency
72-GPU NVLink domain acting as a single massive GPU with NVLink Switch System providing 130 TB/s of low-latency GPU communications.
Multi-GPU Scalability
Scaling Characteristics
Workload Readiness
HPC / Simulation
Designed for AI and high-performance computing (HPC) workloads
Real-Time Serving
Includes a second-generation Transformer Engine enabling FP4 AI for real-time LLM inference
Market Authority
Supercomputer Usage
exascale computer in a single rack
Key Strengths
The GB200 NVL72 excels at handling large-scale AI and machine learning tasks, offering superior performance in model training and inference. Its advanced architecture and high memory bandwidth make it stand out for demanding computational workloads.
Limitations
Potential limitations include high power consumption and cooling requirements. Availability may be constrained by demand and production capacity. Users should ensure compatibility with existing infrastructure and consider the cost implications of deploying such high-performance hardware.
Also in the Lineup
Expert Insight
The GB200 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.