NVIDIA · 2022-03-27

H100

NVL

The NVIDIA H100 NVL variant is optimized for large language model inference, offering up to 5x performance improvement over NVIDIA A100 systems for LLMs up to 70 billion parameters. It features a PCIe form factor, NVLink bridge, and 188GB HBM3 memory for enhanced performance and scalability.

H100 NVL — illustration of the card's form factor
VRAM
94GB
FP32 TFLOPS
60 TFLOPS
CUDA Cores
16,896
TDP
400 W

Provider Marketplace

Cheapest
$1.15/hour
Starting from
Best Value
$2.56/hour
Starting from
Enterprise Choice
$3.58/hour
Starting from

All Cloud Providers

7 Options available
Sharon AI logo
Sharon AICheapest
On-Demand
$1.15/ hour
Estimated Cost
Provision
Hyperstack logo
Reserved · commitment
$1.82/ hour
Estimated Cost
Provision
Nebius logo
Spot · preemptibleThe platform is only available in the `eu-north1` region.
$2.15/ hour
Estimated Cost
Provision
Oblivus logo
Reserved · commitment
$2.56/ hour
Estimated Cost
Provision
RunPod logo
On-Demand
$2.59/ hour
Estimated Cost
Provision
$2.92/ hour
Estimated Cost
Provision
Atlantic.Net logo
Reserved · commitment
$3.58/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP6430 TFLOPS
FP3260 TFLOPS
TF32835 TFLOPS
FP161671 TFLOPS
BF161671 TFLOPS
FP83341 TFLOPS
INT83341 TOPS

Architecture

MicroarchitectureHopper
Tensor Cores4th Generation, Not Published
Matrix EngineTransformer Engine (FP8)
Base Clock1080 MHz
Boost Clock1785 MHz
Transformer EngineYes (FP8)
Sparse AccelerationSupported (unspecified)
Dynamic PrecisionSupported (FP64/TF32/FP32/FP16/FP8/INT8)

Memory & VRAM

Memory TypeHBM3
Total Capacity94GB
Bandwidth3,938 GB/s
Bus Width6,016-bit
ECC SupportEnabled

Connectivity & Scaling

InterconnectNVLink
IB Bandwidth600 GB/s
PCIe InterfaceGen5
TopologyNVLink bridge (point-to-point P2P between adjacent GPUs)
Max GPUs/Node8
Scale-OutNDR Quantum-2 InfiniBand
P2P MemoryPeer-to-peer (P2P) transfers via NVLink bridge

Virtualization

MIG SupportSupported
MIG PartitionsUp to 7 instances
SR-IOVSupported (32 VFs)
vGPU ReadinessSupported (vGPU 16.1 or later: NVIDIA Virtual Compute Server Edition)
GPU SharingMIG

Power & Efficiency

TDP400 W
Peak Power400
PSU RequiredAt least 400 W (via PCIe 16-pin auxiliary power connector, Sense0/Sense1 strapped for 301 W - 450 W power class recommended)
ConnectorsOne PCIe 16-pin auxiliary power connector (12v-2x6 auxiliary power connector)
Thermal LimitsPassive heatsink (requires system airflow); Ambient operating temperature 0°C to 50°C (short term -5°C to 55°C)

Physical Design

Form FactorPCIe
FHFLYes
Slot WidthDual-slot
Weight1,214 grams
CoolingPassive
Rack DensityHigh compute density

Thermals & Cooling

Temp Range0°C to 50°C

Software Ecosystem

CUDACUDA 12.2 or later
Driver StabilityDriver support: Linux: R535 or later; Windows: R535 or later

Server & Deployment

OEM AvailabilityPartner and NVIDIA-Certified Systems with 1–8 GPUs
PreconfiguredPartner and NVIDIA-Certified Systems with 1–8 GPUs

System Compatibility

CPU PairingBridged H100 NVL card pairs should be placed within the same CPU domain (under the same CPU’s topology) for best bridging performance and balanced topology.
NUMAPrefer GPUs bridged under the same CPU or PCIe switch; place same (even) number of GPUs under each CPU socket; maintain balanced CPU:GPU:NIC ratios; GPU counts should be powers of two where possible.
Required PCIePCIe Gen5 (supports Gen5 x16 or Gen5 x8; Gen4 x16 also supported).
MotherboardFull-height, full-length (FHFL) dual-slot PCIe card requiring a PCIe Gen5 (x16/x8) or Gen4 x16 slot; lane and polarity reversal supported; requires system airflow and PCIe 16-pin auxiliary power connector placement/clearance per form factor.
BIOS LimitsUEFI is not supported; SBIOS and OS/hypervisor support must be configured to enable SR-IOV; the card will not boot if the power connector sense indicates less than the default power cap.
OS CompatSupported on Linux and Windows with driver R535 or later; Microsoft Windows 10, Windows 11, Windows Server 2019, and Windows Server 2022 certified; CUDA x86 support from CUDA 12.2 or later; NVIDIA AI Enterprise supported with VMWare.

Benchmarks & Throughput

Structured Sparsity

With sparsity

Transformer Throughput

Transformer Engine with FP8 precision provides up to 4X faster training over the prior generation for GPT-3 (175B) models.

Training Benchmarks

Up to 4X faster training over the prior generation for GPT-3 (175B) models.

Inference Benchmarks

Inference accelerated by up to 30X on the largest models; H100 NVL increases Llama 2 70B performance up to 5x over A100 systems.

Scaling Efficiency

NVLink bridge between two H100 PCIe cards delivers 600 GB/s bidirectional bandwidth (3 bridges / total maximum NVLink bandwidth 600 Gbytes per second).

Multi-GPU Scalability

Scaling Characteristics

ParallelismMIG (up to 7 instances), SR-IOV (32 virtual functions), NVLink peer-to-peer transfers

Workload Readiness

LLM Inference

up to 5x

HPC / Simulation

60 teraFLOPS

Scientific Computing

30 teraFLOPS

Real-Time Serving

up to 5x

Market Authority

Community Benchmarks

$0.09 per million tokens at 66 TPS/user for GPT-OSS-120B using vLLM

Key Strengths

The H100 NVL excels at large-scale AI training and inference tasks, particularly in natural language processing and deep learning models. Its architecture is optimized for transformer models, offering significant performance improvements over previous generations. The GPU's high memory bandwidth and advanced tensor cores make it ideal for demanding computational workloads.

Limitations

The H100 NVL's high power requirements and need for advanced cooling solutions can be a limitation for some deployments. Additionally, its premium pricing and availability constraints may pose challenges for smaller organizations. Users should also consider the infrastructure investment needed to fully leverage its capabilities.

Expert Insight

The H100 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.