NVIDIA

A16

A16 — illustration of the card's form factor
VRAM
16GB
TDP
250 W
Memory
GDDR6
Architecture
Ampere

Provider Marketplace

Cheapest
$0.56/hour
Starting from
Best Value
Awaiting listings
No additional provider yet
Enterprise Choice
Awaiting listings
No additional provider yet

All Cloud Providers

1 Options available
Runcrate logo
RuncrateCheapest
On-Demand
$0.56/ hour
Estimated Cost
View Provider

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Architecture

MicroarchitectureAmpere
Tensor Coresthird-generation Tensor Cores
RT Coressecond-generation RT Cores
Base Clock1312 MHz
Boost Clock1755 MHz

Memory & VRAM

Memory TypeGDDR6
Total Capacity16GB
Bandwidth200 GB/s
Bus Width128-bit
ECC SupportEnabled (by default). Can be disabled via software

Connectivity & Scaling

InterconnectPCI Express
GenerationPCIe Gen4 x16
PCIe InterfaceGen4 xx16

Virtualization

SR-IOVSupported (16 VF per GPU)
vGPU ReadinessSupported (vGPU 13.0 or later; NVIDIA vPC, vApps, vWS, vCS)
GPU SharingvGPU; SR-IOV (16 VF per GPU)
Virt EfficiencyNear bare-metal (vendor claim)

Power & Efficiency

TDP250 W
Peak Power250
ConnectorsCPU 8-pin
Thermal LimitsPassive cooling; requires system airflow; ambient operating temperature 0 °C to 50 °C (short term -5 °C to 55 °C)
EfficiencyNEBS Ready Level 3

Physical Design

Form FactorPCIe
FHFLYes
Slot WidthDual-slot
Weight1088 Grams
CoolingPassive
Rack DensityHigh-density optimized

Thermals & Cooling

Temp Range0 °C to 50 °C
DC Heat250 W

Software Ecosystem

CUDACUDA 11.4 or later
Driver StabilityDriver support R470 or later

Server & Deployment

OEM AvailabilityRefer to qualified servers list
Edge DeployNVIDIA EGX platform
Ref ArchitecturesNVIDIA vPC or NVIDIA RTX vWS (virtual GPU software)

System Compatibility

CPU PairingGen4 x16 connection recommended; Gen3 x16 supported.
Required PCIePCIe Gen4 x16 recommended; Gen3 x16 supported.
MotherboardFull-height, full-length (FHFL) 10.5" dual-slot card; requires x16 PCIe Gen4 connectivity; passive cooling requires system airflow.
Rack Power250 W maximum per board (total board power).
BIOS LimitsUEFI supported; Zero Power not supported; Secure Boot supported (see Root of Trust).
OS CompatMicrosoft Windows 10, Windows Server 2008 R2, Windows Server 2012 R2, and Windows 2016 are supported; Linux supported.

Benchmarks & Throughput

Inference Benchmarks

More than double the encoder throughput versus previous generation M10, providing the multiuser performance required for streaming video and multimedia

Multi-GPU Scalability

Scaling Characteristics

ParallelismSR-IOV support: 16 VF (virtual functions) per GPU

Workload Readiness

Market Authority

Key Strengths

Limitations

Expert Insight

The A16 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.