NVIDIA
H200
SXM

Provider Marketplace
All Cloud Providers
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Compute Performance
Architecture
Memory & VRAM
Connectivity & Scaling
Virtualization
Power & Efficiency
Physical Design
Server & Deployment
System Compatibility
Benchmarks & Throughput
Structured Sparsity
With sparsity.
Inference Benchmarks
Llama2 70B inference 1.9X faster; GPT-3 175B inference 1.6X faster; throughput comparisons vs H100 SXM provided for Llama2 and GPT-3.
Scaling Efficiency
NVIDIA NVLink: 900GB/s
Multi-GPU Scalability
Scaling Characteristics
Workload Readiness
LLM Training
1,979 TFLOPS
LLM Inference
3,958 TFLOPS
Vision Training
1,979 TFLOPS
HPC / Simulation
34 TFLOPS
Scientific Computing
67 TFLOPS
Market Authority
Community Benchmarks
["Llama2 70B Inference: 1.9X Faster","GPT-3 175B Inference: 1.6X Faster","High-Performance Computing: 110X Faster","H200 boosts inference speed by up to 2X compared to H100 GPUs when handling LLMs like Llama2","Throughput comparisons for H200 SXM vs H100 SXM (examples): Llama2 13B and Llama2 70B"]
Key Strengths
Limitations
Also in the Lineup
Expert Insight
The H200 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.