NVIDIA
H200
NVL

Provider Marketplace
All Cloud Providers
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Compute Performance
Architecture
Memory & VRAM
Connectivity & Scaling
Virtualization
Power & Efficiency
Physical Design
Server & Deployment
System Compatibility
Benchmarks & Throughput
Structured Sparsity
With sparsity
Transformer Throughput
The H200 boosts inference speed by up to 2X compared to H100 GPUs when handling LLMs like Llama2.
Inference Benchmarks
NVL: LLM inference can be accelerated up to 1.7x over H100 NVL; H200 generally boosts LLM inference up to 2x vs H100.
Scaling Efficiency
Supports up to 2- or 4-way NVLink bridge (900GB/s per GPU); up to four GPUs can be connected by NVLink for scaling (NVL specific statement).
Multi-GPU Scalability
Scaling Characteristics
Workload Readiness
LLM Inference
Up to 1.7x
HPC / Simulation
Up to 1.3x
Market Authority
Key Strengths
Limitations
Also in the Lineup
Expert Insight
The H200 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.