NVIDIA

RTX 4000 Ada

RTX 4000 Ada — illustration of the card's form factor
VRAM
20GB
FP32 TFLOPS
26.7 TFLOPS
CUDA Cores
6,144
TDP
130 W

Provider Marketplace

Cheapest
$0.46/hour
Starting from
Best Value
$0.76/hour
Starting from
Enterprise Choice
$0.87/hour
Starting from

All Cloud Providers

3 Options available
Paperspace logo
PaperspaceCheapest
Reserved · commitment
$0.46/ hour
Estimated Cost
Provision
DigitalOcean logo
On-Demand
$0.76/ hour
Estimated Cost
Provision
Runcrate logo
On-Demand
$0.87/ hour
Estimated Cost
View Provider

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP3226.7 TFLOPS
FP8327.6 TFLOPS

Architecture

MicroarchitectureAda Lovelace
CUDA Cores6144
Tensor CoresFourth-generation, 192 Tensor Cores
RT CoresThird-generation, 48 RT Cores
Sparse AccelerationSupported
Dynamic PrecisionSupported (FP8)

Memory & VRAM

Memory TypeGDDR6
Total Capacity20GB
Bandwidth360GB/s
Bus Width160-bit
ECC SupportYes (ECC)

Connectivity & Scaling

PCIe InterfacePCIe 4.0 xx16
GPUDirect RDMAYes

Power & Efficiency

TDP130 W
Thermal LimitsActive

Physical Design

Form FactorPCIe
Slot WidthSingle-slot
Dimensions4.4” H x 9.5” L
CoolingActive

Software Ecosystem

CUDACUDA 12.2

Server & Deployment

OEM AvailabilityFind an NVIDIA design and visualization partner or shop the NVIDIA store.

System Compatibility

Required PCIePCIe 4.0 x16
MotherboardPCIe Gen4 x16; single-slot, 4.4” H x 9.5” L
OS CompatWindows 10 and Linux

Benchmarks & Throughput

Structured Sparsity

Effective FP8 teraFLOPS (TFLOPS) using sparsity.

Inference Benchmarks

RTX 4000 accelerates compute-intensive AI workloads, delivering over 1.5X higher inference performance compared to the previous generation

Scaling Efficiency

NVIDIA NVLink No

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksSystem interface: PCIe 4.0 x16; NVLink: No
ParallelismCompute APIs: CUDA 12.2, OpenCL 3.0, DirectCompute; NVIDIA GPUDirect RDMA support; Encode/Decode: 2x encode, 2x decode (+AV1 encode and decode)

Workload Readiness

Diffusion Models

Image generation tested at 512x512 using Stable Diffusion webUI v1.3.1.

HPC / Simulation

significant performance improvements for graphics and simulation workflows on the desktop, such as complex 3D computer-aided design (CAD) and computer-aided engineering (CAE).

Market Authority

Community Benchmarks

["SPECviewperf 2020 geomean test (Graphics)","Arnold v6.0.2 Sol scene (Rendering)","Stable Diffusion webUI v1.3.1 (Generative AI, 512x512 image generation)","TensorRT ResNet-50 V1.5 Inference (Inference, precision: mixed)","NVIDIA Omniverse performance for real-time rendering at 4K with NVIDIA DLSS 3 (Omniverse)"]

Key Strengths

Limitations

Expert Insight

The RTX 4000 Ada represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.