NVIDIA · 2022-03-27
H100
NVL
The NVIDIA H100 NVL variant is optimized for large language model inference, offering up to 5x performance improvement over NVIDIA A100 systems for LLMs up to 70 billion parameters. It features a PCIe form factor, NVLink bridge, and 188GB HBM3 memory for enhanced performance and scalability.

Provider Marketplace
All Cloud Providers
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Compute Performance
Architecture
Memory & VRAM
Connectivity & Scaling
Virtualization
Power & Efficiency
Physical Design
Thermals & Cooling
Software Ecosystem
Server & Deployment
System Compatibility
Benchmarks & Throughput
Structured Sparsity
With sparsity
Transformer Throughput
Transformer Engine with FP8 precision provides up to 4X faster training over the prior generation for GPT-3 (175B) models.
Training Benchmarks
Up to 4X faster training over the prior generation for GPT-3 (175B) models.
Inference Benchmarks
Inference accelerated by up to 30X on the largest models; H100 NVL increases Llama 2 70B performance up to 5x over A100 systems.
Scaling Efficiency
NVLink bridge between two H100 PCIe cards delivers 600 GB/s bidirectional bandwidth (3 bridges / total maximum NVLink bandwidth 600 Gbytes per second).
Multi-GPU Scalability
Scaling Characteristics
Workload Readiness
LLM Inference
up to 5x
HPC / Simulation
60 teraFLOPS
Scientific Computing
30 teraFLOPS
Real-Time Serving
up to 5x
Market Authority
Community Benchmarks
$0.09 per million tokens at 66 TPS/user for GPT-OSS-120B using vLLM
Key Strengths
The H100 NVL excels at large-scale AI training and inference tasks, particularly in natural language processing and deep learning models. Its architecture is optimized for transformer models, offering significant performance improvements over previous generations. The GPU's high memory bandwidth and advanced tensor cores make it ideal for demanding computational workloads.
Limitations
The H100 NVL's high power requirements and need for advanced cooling solutions can be a limitation for some deployments. Additionally, its premium pricing and availability constraints may pose challenges for smaller organizations. Users should also consider the infrastructure investment needed to fully leverage its capabilities.
Also in the Lineup
Expert Insight
The H100 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.