NVIDIA · November 2020

A100

80GB SXM

The NVIDIA A100 80GB SXM is a high-performance GPU designed for data centers, targeting AI, machine learning, and high-performance computing workloads. It is part of the Ampere architecture, offering significant improvements in memory capacity and bandwidth over its predecessors. The 80GB variant provides enhanced memory for large-scale models and datasets, making it ideal for demanding applications.

A100 80GB SXM — illustration of the card's form factor
VRAM
80GB
CUDA Cores
6,912
TDP
400 W
Memory
HBM2e

Provider Marketplace

Cheapest
$0.60/hour
Starting from
Best Value
$0.98/hour
Starting from
Enterprise Choice
$2.79/hour
Starting from

All Cloud Providers

9 Options available
Vast.ai logo
Vast.aiCheapest
On-DemandCroatia, HR
$0.60/ hour
Estimated Cost
Provision
CloudRift logo
Reserved · commitment
$0.89/ hour
Estimated Cost
Provision
Verda logo
Spot · preemptible
$0.90/ hour
Estimated Cost
Provision
Oblivus logo
Reserved · commitment
$0.95/ hour
Estimated Cost
Provision
Hyperstack logo
Reserved · commitment
$0.98/ hour
Estimated Cost
Provision
$1.38/ hour
Estimated Cost
Provision
RunPod logo
On-Demand
$1.39/ hour
Estimated Cost
Provision
TensorDock logo
On-Demand
$1.80/ hour
Estimated Cost
Provision
Lambda Labs logo
On-Demand
$2.79/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

TF32312 TFLOPS*
FP16624 TFLOPS*
BF16624 TFLOPS*
INT81248 TOPS*

Architecture

MicroarchitectureAmpere
Tensor CoresThird-generation Tensor Cores
Matrix EngineTensor Cores (TF32 / BFLOAT16 / FP16)
Sparse AccelerationSupported (structural sparsity, up to 2X performance)
Dynamic PrecisionSupported (FP32 to INT4)

Memory & VRAM

Memory TypeHBM2e
Total Capacity80GB
Bandwidth2,039 GB/s
Unified MemoryYes (unified memory per node: up to 1.3 TB)

Connectivity & Scaling

InterconnectNVLink
Generation3rd-generation NVLink
IB Bandwidth600 GB/s
PCIe InterfacePCIe Gen4
TopologyNVSwitch / NVLink (up to 16 GPUs interconnected)
Max GPUs/Node16
Scale-OutNVIDIA InfiniBand

Virtualization

MIG SupportSupported
MIG Partitions7 instances
K8s ReadinessSupported (MIG works with Kubernetes)
GPU SharingMIG

Power & Efficiency

TDP400 W

Physical Design

Form FactorSXM

Thermals & Cooling

DC Heat400 W

Software Ecosystem

PyTorchsupported
TensorFlowsupported

Server & Deployment

OEM AvailabilityNVIDIA HGX™ A100-Partner and NVIDIA-Certified Systems with 4,8, or 16 GPUs NVIDIA DGX™ A100 with 8 GPUs
PreconfiguredNVIDIA-Certified System
DGX/HGXNVIDIA HGX™ A100-Partner and NVIDIA-Certified Systems with 4,8, or 16 GPUs NVIDIA DGX™ A100 with 8 GPUs

System Compatibility

Required PCIePCIe Gen4
MotherboardHGX A100 server boards
OS CompatKubernetes; containers; hypervisor-based server virtualization; VMware vSphere

Benchmarks & Throughput

Structured Sparsity

Supported (up to 2X vs dense)

Training Benchmarks

A100 80GB reaches up to 1.3 TB of unified memory per node and delivers up to a 3X throughput increase over A100 40GB (DLRM).

Inference Benchmarks

On state-of-the-art conversational AI models like BERT, A100 accelerates inference throughput up to 249X over CPUs.

Scaling Efficiency

NVLink/NVSwitch/InfiniBand scaling: possible to scale to thousands of A100 GPUs; NVLink/NVSwitch interconnect up to 600 GB/s.

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksBut scale-out solutions are often bogged down by datasets scattered across multiple servers.
ParallelismAn A100 GPU can be partitioned into as many as seven GPU instances, fully isolated at the hardware level with their own high-bandwidth memory, cache, and compute cores.

Workload Readiness

LLM Training

312 TFLOPS (TF32 peak for A100 80GB SXM)

LLM Inference

1248 TOPS (INT8 Tensor Core peak for A100 80GB SXM with sparsity)

Vision Training

624 TFLOPS (FP16 Tensor Core peak for A100 80GB SXM with sparsity)

HPC / Simulation

9.7 TFLOPS (FP64 peak for NVIDIA A100 for NVLink / SXM)

Scientific Computing

19.5 TFLOPS (FP64 Tensor Core peak)

Real-Time Serving

Up to 7 MIGs @ 10GB (Multi-Instance GPU for A100 80GB SXM)

Market Authority

Key Strengths

This GPU excels at AI training and inference, offering exceptional performance for deep learning frameworks. Its large memory capacity and high bandwidth make it particularly effective for large-scale models and data-intensive tasks. The A100's support for multi-instance GPU (MIG) technology allows for efficient resource partitioning, enhancing its versatility.

Limitations

While the A100 80GB SXM offers exceptional performance, its high power consumption and cooling requirements may limit its use to well-equipped data centers. The SXM form factor restricts compatibility to specific platforms, and its premium pricing can be a barrier for smaller organizations. Availability may also be constrained by high demand and production limitations.

Expert Insight

The A100 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.