NVIDIA

L40

L40 — illustration of the card's form factor
VRAM
48GB
FP32 TFLOPS
90.5 TFLOPS
CUDA Cores
18,176
TDP
300 W

Provider Marketplace

Cheapest
$0.44/hour
Starting from
Best Value
$0.79/hour
Starting from
Enterprise Choice
$0.97/hour
Starting from

All Cloud Providers

8 Options available
Oblivus logo
OblivusCheapest
Reserved · commitment
$0.44/ hour
Estimated Cost
Provision
RunPod logo
On-Demand
$0.69/ hour
Estimated Cost
Provision
Hyperstack logo
Reserved · commitment
$0.70/ hour
Estimated Cost
Provision
CoreWeave logo
Spot · preemptibleNORTH AMERICA
$0.78/ hour
Estimated Cost
Provision
$0.79/ hour
Estimated Cost
Provision
$0.86/ hour
Estimated Cost
Provision
TensorDock logo
On-Demand
$0.95/ hour
Estimated Cost
Provision
Runcrate logo
On-Demand
$0.97/ hour
Estimated Cost
View Provider

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP3290.5 TFLOPS
TF3290.5 TFLOPS
FP16181.05 TFLOPS
BF16181.05 TFLOPS
FP8362 TFLOPS
INT8362 TOPS
INT4724 TOPS

Architecture

MicroarchitectureAda Lovelace
CUDA Cores18176
Tensor CoresFourth Generation, 568 Tensor Cores
RT CoresThird Generation, 142 RT Cores
Sparse AccelerationSupported (structural sparsity)
Dynamic PrecisionSupported (FP8/FP16/BF16/TF32/FP32/INT8/INT4)

Memory & VRAM

Memory TypeGDDR6
Total Capacity48GB
Bandwidth864GB/s
ECC SupportYes (ECC)
Memory PoolingNot Supported

Connectivity & Scaling

InterconnectPCIe
GenerationPCIe Gen4
IB Bandwidth64GB/s bi-directional
PCIe InterfacePCIe Gen4 xx16

Virtualization

MIG SupportNot Supported
vGPU ReadinessSupported (NVIDIA vPC/vApps, NVIDIA RTX Virtual Workstation (vWS))
GPU SharingvGPU
Virt EfficiencyNear bare-metal (vendor claim)

Power & Efficiency

TDP300 W
Peak Power300
Connectors1x PCIe CEM5 16-pin
Thermal LimitsPassive
EfficiencyNEBS Level 3

Physical Design

Form FactorPCIe
FHFLYes
Slot WidthDual-slot
CoolingPassive

Software Ecosystem

Driver Stabilityenterprise-class stability and reliability

Server & Deployment

OEM AvailabilityPackaged in a dual-slot, passively cooled and power-efficient design, the L40 is available in a wide variety of NVIDIA-Certified Systems ™ from leading OEM vendors.
PreconfiguredNVIDIA OVX systems

System Compatibility

Required PCIePCIe Gen4x16
MotherboardFull-height, full-length (FHFL) dual-slot
Rack Power300W

Benchmarks & Throughput

Structured Sparsity

Supported

Inference Benchmarks

delivering 5X higher inference performance compared to the previous generation

Scaling Efficiency

NVLink Support: No

Multi-GPU Scalability

Scaling Characteristics

Network BottlenecksInterconnect: PCIe Gen4x16 (64GB/s bi-directional); NVLink: not supported
ParallelismvGPU software support: Yes; vGPU profiles supported: 1; MIG: No; NVLink: No

Workload Readiness

Vision Training

Supported for visual computing workloads

Diffusion Models

Accelerates image generative AI / inference

HPC / Simulation

Supported for large-scale modeling and simulation

Scientific Computing

90.5

Market Authority

Key Strengths

Limitations

Expert Insight

The L40 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.