NVIDIA
A40

Provider Marketplace
All Cloud Providers
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Compute Performance
Architecture
Memory & VRAM
Connectivity & Scaling
Virtualization
Power & Efficiency
Physical Design
Software Ecosystem
Server & Deployment
System Compatibility
Benchmarks & Throughput
Structured Sparsity
Hardware support for structural sparsity doubles the throughput for inferencing.
Training Benchmarks
Up to 3X Faster AI Training Performance (BERT pre-training throughput)
Inference Benchmarks
Hardware support for structural sparsity doubles the throughput for inferencing.
Scaling Efficiency
Connect two A40 GPUs together to scale from 48GB of GPU memory to 96GB.
Multi-GPU Scalability
Scaling Efficiency
Scaling Characteristics
Workload Readiness
LLM Training
Up to 3X Faster AI Training Performance
LLM Inference
Hardware support for structural sparsity provides up to double the throughput for inferencing.
Vision Training
Tensor Float 32 (TF32) precision provides up to 5X the training throughput over the previous generation to accelerate AI and data science model training without requiring any code changes.
HPC / Simulation
Up to 50% Faster Single Precision (FP32) HPC Performance
Scientific Computing
Peak FP32 TFLOPS (non-Tensor) 37.4
Market Authority
Community Benchmarks
SPECviewperf 2020; Iray 2020.1; NAMD; BERT pre-training throughput
Key Strengths
Limitations
Also in the Lineup
Expert Insight
The A40 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.