NVIDIA

A40

A40 — illustration of the card's form factor
VRAM
48GB
FP32 TFLOPS
37.4 TFLOPS
CUDA Cores
10,752
TDP
300 W

Provider Marketplace

Cheapest
$0.35/hour
Starting from
Best Value
$2.05/hour
Starting from
Enterprise Choice
$2.11/hour
Starting from

All Cloud Providers

4 Options available
RunPod logo
RunPodCheapest
On-Demand
$0.35/ hour
Estimated Cost
Provision
E2E Networks logo
On-Demand
$1.44/ hour
Estimated Cost
Provision
Runcrate logo
On-Demand
$2.05/ hour
Estimated Cost
View Provider
$2.11/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP3237.4 TFLOPS
TF3274.8 | 149.6* TFLOPS
FP16149.7 | 299.4* TFLOPS
BF16149.7 | 299.4* TFLOPS
INT8299.3 | 598.6* TOPS
INT4598.7 | 1,197.4* TOPS

Architecture

MicroarchitectureAmpere
CUDA Cores10752
Tensor CoresThird-Generation, 336 Tensor Cores
RT CoresSecond-Generation, 84 RT Cores
Matrix EngineTensor Float 32 (TF32)
Sparse AccelerationSupported (structural sparsity)
Dynamic PrecisionSupported (TF32/FP16/BF16/INT8/INT4)

Memory & VRAM

Memory TypeGDDR6
Total Capacity48GB
Bandwidth696 GB/s
ECC SupportYes (ECC)
Unified MemoryYes (NVLink single scalable memory)
NUMA AwarenessNo (MIG support)
Memory PoolingYes (NVLink memory scaling to 96GB across two GPUs)

Connectivity & Scaling

InterconnectNVIDIA NVLink
GenerationThird-Generation NVIDIA NVLink
IB Bandwidth112.5 GB/s (bidirectional)
PCIe InterfacePCIe Gen4
Topology2-way NVLink (low profile, 2-slot)
P2P MemoryYes

Virtualization

MIG SupportNot Supported
vGPU ReadinessSupported (NVIDIA vGPU)
GPU SharingvGPU (NVIDIA vGPU software)

Power & Efficiency

TDP300 W
Peak Power300
Connectors8-pin CPU
Thermal LimitsPassive
EfficiencyNEBS Ready Level 3

Physical Design

Form Factor4.4" (H) x 10.5" (L) Dual Slot
Slot WidthDual Slot
Dimensions4.4" (H) x 10.5" (L)
CoolingPassive

Software Ecosystem

CUDA11.1.74
PyTorchPyTorch (2/3) Phase 1 and (1/3) Phase 2
Compiler StackCUDA, DirectCompute, OpenCL, OpenACC
Driver StabilityEnterprise drivers with extensive testing and ISV certifications to ensure optimal performance and stability

Server & Deployment

OEM AvailabilityValidated with NVIDIA-Certified systems from worldwide OEMs
PreconfiguredAvailable via NVIDIA-Certified systems from worldwide OEMs
Edge DeployNVIDIA EGX platform with NVIDIA vGPU software

System Compatibility

Required PCIePCI Express Gen 4
Motherboardvalidated with a wide range of NVIDIA-Certified systems from worldwide OEMs
Rack Power300 W
OS CompatWindows 10 and Linux

Benchmarks & Throughput

Structured Sparsity

Hardware support for structural sparsity doubles the throughput for inferencing.

Training Benchmarks

Up to 3X Faster AI Training Performance (BERT pre-training throughput)

Inference Benchmarks

Hardware support for structural sparsity doubles the throughput for inferencing.

Scaling Efficiency

Connect two A40 GPUs together to scale from 48GB of GPU memory to 96GB.

Multi-GPU Scalability

Scaling Efficiency

Single GPUup to 2X as power efficient as the previous generation

Scaling Characteristics

Network BottlenecksNVIDIA NVLink 112.5 GB/s (bidirectional); PCIe Gen4: 64GB/s
ParallelismvGPU software support; NVLink 2-way; MIG not supported

Workload Readiness

LLM Training

Up to 3X Faster AI Training Performance

LLM Inference

Hardware support for structural sparsity provides up to double the throughput for inferencing.

Vision Training

Tensor Float 32 (TF32) precision provides up to 5X the training throughput over the previous generation to accelerate AI and data science model training without requiring any code changes.

HPC / Simulation

Up to 50% Faster Single Precision (FP32) HPC Performance

Scientific Computing

Peak FP32 TFLOPS (non-Tensor) 37.4

Market Authority

Community Benchmarks

SPECviewperf 2020; Iray 2020.1; NAMD; BERT pre-training throughput

Key Strengths

Limitations

Expert Insight

The A40 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.