NVIDIA

H200

NVL

H200 NVL — illustration of the card's form factor
VRAM
141GB
FP32 TFLOPS
60 TFLOPS
TDP
600 W
Memory
HBM3E

Provider Marketplace

Cheapest
$3.62/hour
Starting from
Best Value
Awaiting listings
No additional provider yet
Enterprise Choice
Awaiting listings
No additional provider yet

All Cloud Providers

1 Options available
Massed Compute logo
On-Demand
$3.62/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP6430 TFLOPS
FP3260 TFLOPS
TF32835 TFLOPS
FP161,671 TFLOPS
BF161,671 TFLOPS
FP83,341 TFLOPS
INT83,341 TFLOPS TOPS

Architecture

MicroarchitectureHopper
Sparse AccelerationSupported (With sparsity)
Dynamic PrecisionSupported (FP64, FP64 Tensor Core, FP32, TF32 Tensor Core, BFLOAT16 Tensor Core, FP16 Tensor Core, FP8 Tensor Core, INT8 Tensor Core)

Memory & VRAM

Memory TypeHBM3E
Total Capacity141GB
Bandwidth4.8TB/s

Connectivity & Scaling

InterconnectNVIDIA NVLink bridge; PCIe Gen5
GenerationPCIe Gen5
IB Bandwidth900GB/s per GPU; PCIe Gen5: 128GB/s
PCIe InterfacePCIe Gen5
Topology2- or 4-way NVIDIA NVLink bridge
Max GPUs/Node8

Virtualization

MIG SupportSupported
MIG PartitionsUp to 7 MIGs @16.5GB each
GPU SharingMIG

Power & Efficiency

TDP600 W
Thermal LimitsPCIe Dual-slot air-cooled

Physical Design

Form FactorPCIe
Slot WidthDual-slot
Coolingair-cooled
Rack Densitylower-power, air-cooled enterprise rack designs

Server & Deployment

OEM AvailabilityNVIDIA MGX™ H200 NVL partner and NVIDIA-Certified Systems with up to 8 GPUs
PreconfiguredNVIDIA MGX™ H200 NVL partner and NVIDIA-Certified Systems with up to 8 GPUs
Rack-ScaleNVIDIA H200 NVL is ideal for lower-power, air-cooled enterprise rack designs that require flexible configurations, delivering acceleration for every AI and HPC workload regardless of size.

System Compatibility

Required PCIePCIe Gen5
MotherboardPCIe Dual-slot air-cooled
Rack PowerUp to 600W (configurable)

Benchmarks & Throughput

Structured Sparsity

With sparsity

Transformer Throughput

The H200 boosts inference speed by up to 2X compared to H100 GPUs when handling LLMs like Llama2.

Inference Benchmarks

NVL: LLM inference can be accelerated up to 1.7x over H100 NVL; H200 generally boosts LLM inference up to 2x vs H100.

Scaling Efficiency

Supports up to 2- or 4-way NVLink bridge (900GB/s per GPU); up to four GPUs can be connected by NVLink for scaling (NVL specific statement).

Multi-GPU Scalability

Scaling Characteristics

Network Bottlenecks2- or 4-way NVIDIA NVLink bridge: 900GB/s per GPU PCIe Gen5: 128GB/s
ParallelismUp to 7 MIGs @16.5GB each

Workload Readiness

LLM Inference

Up to 1.7x

HPC / Simulation

Up to 1.3x

Market Authority

Key Strengths

Limitations

Expert Insight

The H200 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.