NVIDIA

RTX A4000

RTX A4000 — illustration of the card's form factor
VRAM
16GB
FP32 TFLOPS
19.2 TFLOPS
CUDA Cores
6,144
TDP
140 W

Provider Marketplace

Cheapest
$0.10/hour
Starting from
Best Value
$0.67/hour
Starting from
Enterprise Choice
$0.88/hour
Starting from

All Cloud Providers

4 Options available
TensorDock logo
TensorDockCheapest
On-Demand
$0.10/ hour
Estimated Cost
Provision
Hyperstack logo
Reserved · commitment
$0.11/ hour
Estimated Cost
Provision
Paperspace logo
Reserved · commitment
$0.67/ hour
Estimated Cost
Provision
Runcrate logo
On-Demand
$0.88/ hour
Estimated Cost
View Provider

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP3219.2 TFLOPS

Architecture

MicroarchitectureAmpere
CUDA Cores6144
Tensor CoresThird-generation, 192 Tensor Cores
RT CoresSecond-generation, 48 RT Cores
Sparse AccelerationSupported (structural sparsity)

Memory & VRAM

Memory TypeGDDR6
Total Capacity16GB
Bandwidth448 GB/s
Bus Width256-bit
ECC SupportYes

Connectivity & Scaling

PCIe InterfaceGen 4 xx16

Power & Efficiency

TDP140 W
Peak Power140
Connectors1x 6-pin PCIe
Thermal LimitsActive

Physical Design

Form FactorPCIe
Slot WidthSingle-slot
Dimensions4.4” H x 9.5” L
CoolingActive

Software Ecosystem

CUDACUDA 11.6

System Compatibility

Required PCIePCIe 4.0 x16
MotherboardSingle-slot PCIe form factor (fits into a wide range of workstation chassis)
OS CompatWindows 10, Windows 11, and Linux

Benchmarks & Throughput

Structured Sparsity

Supported (hardware-support for structural sparsity)

Training Benchmarks

Up to 11X faster training performance compared to the previous generation

Multi-GPU Scalability

Scaling Characteristics

ParallelismCompute APIs: CUDA 11.6, DirectCompute, OpenCL 3.0

Workload Readiness

LLM Training

Boost AI and data science model training with up to 11X faster training performance compared to the previous generation with hardware-support for structural sparsity.

LLM Inference

AI-accelerated compute

Market Authority

Enterprise Cases

Design Showcase: RTX All Stars

Key Strengths

Limitations

Expert Insight

The RTX A4000 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.