NVIDIA · August 2023

L40S

The NVIDIA L40S is a high-performance GPU designed for datacenter environments, targeting AI workloads, graphics rendering, and virtualization. It is part of the Ada Lovelace architecture, offering enhanced performance and efficiency over previous generations. The L40S is tailored for enterprise applications, providing robust support for AI and graphics-intensive tasks.

L40S — illustration of the card's form factor
VRAM
48GB
FP32 TFLOPS
91.6 TFLOPS
CUDA Cores
18,176
TDP
350 W

Provider Marketplace

Cheapest
$0.54/hour
Starting from
Best Value
$1.50/hour
Starting from
Enterprise Choice
$4.43/hour
Starting from

All Cloud Providers

18 Options available
CloudRift logo
CloudRiftCheapest
Reserved · commitment
$0.54/ hour
Estimated Cost
Provision
Verda logo
Spot · preemptible
$0.69/ hour
Estimated Cost
Provision
Nebius logo
Spot · preemptible
$0.74/ hour
Estimated Cost
Provision
RunPod logo
On-Demand
$0.79/ hour
Estimated Cost
Provision
$0.88/ hour
Estimated Cost
Provision
Runcrate logo
On-Demand
$0.97/ hour
Estimated Cost
View Provider
CoreWeave logo
Spot · preemptibleNORTH AMERICA
$0.98/ hour
Estimated Cost
Provision
RedSwitches logo
On-DemandMontrealCanada
$1.07/ hour
Estimated Cost
Provision
E2E Networks logo
On-Demand
$1.20/ hour
Estimated Cost
Provision
Crusoe Cloud logo
On-Demand
$1.50/ hour
Estimated Cost
Provision
DigitalOcean logo
On-Demand
$1.57/ hour
Estimated Cost
Provision
Atlantic.Net logo
Reserved · commitment
$1.58/ hour
Estimated Cost
Provision
OVHcloud logo
On-Demand
$1.80/ hour
Estimated Cost
Provision
Modal logo
On-Demand
$1.95/ hour
Estimated Cost
Provision
Lightning AI logo
On-Demand
$2.14/ hour
Estimated Cost
Provision
$2.25/ hour
Estimated Cost
Provision
Replicate logo
On-Demand
$3.51/ hour
Estimated Cost
Provision
IBM Cloud logo
On-Demand
$4.43/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP3291.6 TFLOPS
TF32183 TFLOPS
FP16362.05 TFLOPS
BF16362.05 TFLOPS
FP8733 TFLOPS
INT8733 TOPS
INT4733 TOPS

Architecture

MicroarchitectureAda Lovelace
CUDA Cores18176
Tensor CoresFourth-Generation, 568 Tensor Cores
RT CoresThird-Generation, 142 RT Cores
Matrix EngineTransformer Engine (FP8/FP16)
Transformer EngineYes (Transformer Engine)
Sparse AccelerationSupported (structural sparsity)
Dynamic PrecisionSupported (FP8/FP16/BF16/TF32)

Memory & VRAM

Memory TypeGDDR6
Total Capacity48GB
Bandwidth864GB/s
ECC SupportYes (ECC)
Memory PoolingNot Supported

Connectivity & Scaling

InterconnectPCIe
GenerationPCIe Gen4
IB Bandwidth64GB/s bidirectional
PCIe InterfaceGen4 xx16
Scale-OutEthernet

Virtualization

MIG SupportNot Supported
vGPU ReadinessSupported
GPU SharingvGPU
Virt EfficiencyNear bare-metal (vendor claim)

Power & Efficiency

TDP350 W
Peak Power350
Connectors16-pin
Thermal LimitsPassive
EfficiencyNEBS Level 3

Physical Design

Slot Widthdual slot
Dimensions4.4" (H) x 10.5" (L)
CoolingPassive

Software Ecosystem

Driver StabilityL40S GPU is optimized for 24/7 enterprise data center operations and designed, built, tested, and supported by NVIDIA to ensure maximum performance, durability, and uptime.

Server & Deployment

OEM AvailabilityDell, Hewlett Packard Enterprise, Lenovo, Supermicro, and others
PreconfiguredNVIDIA OVX™ Servers
Rack-ScaleNVIDIA OVX L40S (scalable data center infrastructure)

System Compatibility

Required PCIePCIe Gen4 x16
MotherboardRequires a PCIe Gen4 x16 slot; dual-slot card, 4.4" (H) x 10.5" (L); 16-pin power connector

Benchmarks & Throughput

Structured Sparsity

Supported

Transformer Throughput

Transformer Engine dramatically accelerates AI performance and improves memory utilization for both training and inference.

Scaling Efficiency

No NVLink (PCIe Gen4 x16 interconnect)

Multi-GPU Scalability

Scaling Characteristics

ParallelismVirtual GPU (vGPU) Software Support: Yes; Multi-Instance GPU (MIG) Support: No; NVIDIA® NVLink® Support: No

Workload Readiness

LLM Training

Supported

LLM Inference

Supported

Diffusion Models

Supported

Multimodal AI

Supported

HPC / Simulation

Supported

Scientific Computing

Supported

Market Authority

Key Strengths

The L40S excels in AI training and inference, offering significant performance improvements for deep learning models. It is also highly effective for graphics rendering and virtualization, making it a versatile choice for mixed workloads in datacenters.

Limitations

While the L40S offers impressive performance, it may come at a higher cost compared to other GPUs in its class. Availability might be limited due to high demand, and users should ensure their systems can accommodate its power and cooling requirements.

Expert Insight

The L40S represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.