NVIDIA · Q2 2023

GB300

NVL72

The NVIDIA GB300 NVL72 is a high-performance GPU designed for datacenter applications, particularly in AI and HPC workloads. It is part of NVIDIA's latest architecture, offering significant improvements in performance and efficiency. The NVL72 variant is optimized for multi-GPU configurations, making it ideal for large-scale AI training and inference tasks.

GB300 NVL72 — illustration of the card's form factor
Memory
HBM3E
Architecture
Blackwell
Form Factor
rack-scale architecture

Provider Marketplace

Cheapest
$4.25/hour
Starting from
Best Value
Awaiting listings
No additional provider yet
Enterprise Choice
$4.31/hour
Starting from

All Cloud Providers

2 Options available
Runcrate logo
RuncrateCheapest
On-Demand
$4.25/ hour
Estimated Cost
Provision
Verda logo
Spot · preemptible
$4.31/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Architecture

MicroarchitectureBlackwell
Sparse AccelerationSupported
Dynamic PrecisionSupported (FP4/FP8/FP6/FP16/BF16/TF32/FP32/FP64/INT8)

Memory & VRAM

Memory TypeHBM3E

Connectivity & Scaling

InterconnectNVLink
GenerationFifth-Generation NVIDIA NVLink
Topology72-GPU NVLink domain
Scale-OutNVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet
GPUDirect RDMAYes

Power & Efficiency

Thermal Limitsfully liquid-cooled, rack-scale architecture

Physical Design

Form Factorrack-scale architecture
Coolingfully liquid-cooled
Rack Densityrack-scale architecture

Thermals & Cooling

Liquid Coolingfully liquid-cooled

Server & Deployment

OEM AvailabilityAvailable Now
Rack-Scalefully liquid-cooled, rack-scale architecture

Multi-GPU Scalability

Scaling Characteristics

Parallelism72-GPU NVLink domain using fifth-generation NVLink scale-up interconnect with NVLink Switches and 130 TB/s NVLink bandwidth

Workload Readiness

LLM Inference

It’s purpose-built for test-time scaling inference and AI reasoning tasks.

Diffusion Models

GB300 NVL72 introduces cutting-edge capabilities for diffusion-based video generation models.

Real-Time Serving

It’s purpose-built for test-time scaling inference and AI reasoning tasks.

Market Authority

Enterprise Cases

["Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems","NVIDIA DGX GB300 Comes Online at Naval Postgraduate School"]

Key Strengths

The GB300 NVL72 excels in AI training and inference, offering superior performance for deep learning models. Its architecture is optimized for high throughput and low latency, making it a top choice for scientific computing and complex simulations. The NVL72's multi-GPU capabilities enhance its performance in parallel processing tasks.

Limitations

The GB300 NVL72's high power consumption and cooling requirements may limit its use in environments with restricted power or cooling capabilities. Its availability might be constrained due to high demand and production limitations. Users should ensure compatibility with existing infrastructure to fully leverage its capabilities.

Expert Insight

The GB300 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.