NVIDIA · 2022-03-27

H100

SXM

The NVIDIA H100 SXM variant features exceptional performance and scalability for a wide range of workloads. It includes fourth-generation Tensor Cores and a Transformer Engine with FP8 precision, providing up to 4X faster training over the prior generation for large language models.

H100 SXM — illustration of the card's form factor
VRAM
80GB
FP32 TFLOPS
67 TFLOPS
CUDA Cores
16,896
TDP
700 W

Provider Marketplace

Cheapest
$1.13/hour
Starting from
Best Value
$2.69/hour
Starting from
Enterprise Choice
$5.50/hour
Starting from

All Cloud Providers

20 Options available
Lightning AI logo
Lightning AICheapest
On-Demand
$1.13/ hour
Estimated Cost
Provision
Runcrate logo
On-Demand
$1.35/ hour
Estimated Cost
Provision
Genesis Cloud logo
Reserved · commitment
$1.60/ hour
Estimated Cost
Provision
Verda logo
Spot · preemptible
$1.63/ hour
Estimated Cost
Provision
TensorDock logo
On-Demand
$2.25/ hour
Estimated Cost
Provision
Vultr logo
On-Demand
$2.30/ hour
Estimated Cost
Provision
CoreWeave logo
Spot · preemptibleEUROPE
$2.44/ hour
Estimated Cost
Provision
Civo logo
Reserved · commitment
$2.49/ hour
Estimated Cost
Provision
Taiga Cloud logo
On-Demand
$2.50/ hour
Estimated Cost
Provision
RunPod logo
On-Demand
$2.69/ hour
Estimated Cost
Provision
Jarvis Labs logo
On-Demand
$2.69/ hour
Estimated Cost
Provision
Oblivus logo
Reserved · commitment
$2.71/ hour
Estimated Cost
Provision
Hyperstack logo
Reserved · commitment
$2.72/ hour
Estimated Cost
Provision
$2.89/ hour
Estimated Cost
Provision
Together logo
Reserved · commitment
$3.19/ hour
Estimated Cost
Provision
DigitalOcean logo
Reserved · commitment
$3.26/ hour
Estimated Cost
Provision
Modal logo
On-Demand
$3.95/ hour
Estimated Cost
Provision
Lambda Labs logo
On-Demand
$3.99/ hour
Estimated Cost
Provision
$4.99/ hour
Estimated Cost
Provision
Crusoe Cloud logo
On-Demand
$5.50/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP6434 TFLOPS
FP3267 TFLOPS
TF32989 TFLOPS
FP161979 TFLOPS
BF161979 TFLOPS
FP83958 TFLOPS
INT83958 TOPS

Architecture

MicroarchitectureHopper
Tensor Coresfourth-generation Tensor Cores
Matrix EngineTransformer Engine (FP8)
Transformer EngineYes (Transformer Engine with FP8)
Sparse AccelerationSupported (with sparsity; format not specified)
Dynamic PrecisionSupported (FP64, TF32, FP32, FP16, INT8, FP8)

Memory & VRAM

Total Capacity80GB
Bandwidth3.35TB/s

Connectivity & Scaling

InterconnectNVIDIA NVLink
Generationfourth-generation NVLink
IB Bandwidth900 GB/s
PCIe InterfacePCIe Gen5
TopologyNVLink and NVSwitch
Max GPUs/Node8
Scale-OutNDR Quantum-2 InfiniBand
P2P MemoryYes

Virtualization

MIG SupportSupported
MIG PartitionsUp to 7 MIGS @ 10GB each
GPU SharingMIG

Power & Efficiency

TDP700 W

Physical Design

Form FactorSXM

Thermals & Cooling

ThrottlingHardware slowdown (50% clock slowdown) at GPU TLIMIT = -2°C; hardware shutdown at GPU TLIMIT = -5°C.
DC HeatUp to 700W (configurable)

Server & Deployment

OEM AvailabilityNVIDIA HGX H100 Partner and NVIDIA- Certified Systems™ with 4 or 8 GPUs NVIDIA DGX H100 with 8 GPUs
PreconfiguredNVIDIA HGX H100 Partner and NVIDIA- Certified Systems™ with 4 or 8 GPUs NVIDIA DGX H100 with 8 GPUs
DGX/HGXNVIDIA HGX H100 Partner and NVIDIA- Certified Systems™ with 4 or 8 GPUs NVIDIA DGX H100 with 8 GPUs
Ref ArchitecturesNVIDIA Grace Hopper CPU+GPU architecture; NVIDIA HGX H100

System Compatibility

Required PCIePCIe Gen5

Benchmarks & Throughput

Structured Sparsity

With sparsity

Scaling Efficiency

NVIDIA NVLink: 900GB/s

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksFourth-generation NVLink (900 GB/s), NDR Quantum-2 InfiniBand, PCIe Gen5, and NVIDIA Magnum IO are presented as the communications stack to accelerate GPU communication across nodes.
ParallelismMulti-Instance GPU (MIG) support: Up to 7 MIGs @ 10GB each.

Workload Readiness

LLM Training

3,958 teraFLOPS

LLM Inference

3,958 TOPS

Vision Training

1,979 teraFLOPS

Diffusion Models

1,979 teraFLOPS

Multimodal AI

80GB

Reinforcement Learning

989 teraFLOPS

HPC / Simulation

34 teraFLOPS

Scientific Computing

3.35TB/s

Real-Time Serving

3,958 TOPS

Market Authority

Key Strengths

The H100 SXM excels in AI and machine learning workloads, particularly in training large neural networks and performing inference at scale. It offers significant performance improvements over its predecessors due to its advanced architecture and increased memory bandwidth. The H100 is also well-suited for high-performance computing (HPC) applications, providing exceptional computational power and efficiency.

Limitations

One limitation of the H100 SXM is its high power consumption, which may not be suitable for all datacenter environments. Additionally, its reliance on specific server platforms and cooling solutions can limit deployment flexibility. Availability can be constrained due to high demand and production capacities, potentially leading to longer lead times for procurement.

Expert Insight

The H100 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.