NVIDIA

P4

P4 — illustration of the card's form factor
VRAM
8 GB
FP32 TFLOPS
5.5 TFLOPS
CUDA Cores
2,560
Memory
GDDR5

Provider Marketplace

Cheapest
$0.15/hour
Starting from
Best Value
Awaiting listings
No additional provider yet
Enterprise Choice
$0.27/hour
Starting from

All Cloud Providers

2 Options available
Alibaba Cloud logo
On-DemandChina (Beijing),China (Shanghai),China (Hangzhou), Singapore
$0.15/ hour
Estimated Cost
Provision
Google Cloud logo
Reserved · commitment
$0.27/ hour
Estimated Cost
Provision

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Compute Performance

FP325.5 TFLOPS
INT822 TOPS

Architecture

MicroarchitecturePascal
CUDA Cores2560
Base Clock885 MHz
Boost Clock1531 MHz

Memory & VRAM

Memory TypeGDDR5
Total Capacity8 GB
Bandwidth192 GB/s
Bus Width256-bit
ECC SupportSupported (Enabled by default)

Connectivity & Scaling

InterconnectPCI Express
GenerationPCI Express Gen3
PCIe InterfacePCI Express 3.0 x×16

Power & Efficiency

Thermal LimitsPassively cooled; requires system air flow to operate within thermal limits; Operating temperature 0 °C to 55 °C.

Physical Design

Form FactorPCIe
Slot WidthSingle-slot
Dimensions68.58 mm × 167.64 mm
Weight240 Grams
CoolingPassive
Rack DensityDensity-optimized, scale-out servers

Thermals & Cooling

Temp Range0 °C to 55 °C
ThrottlingGPU Boost dynamically adjusts the GPU clock to maximize performance within thermal limits.

Software Ecosystem

Kernel OptimPage Migration Engine
Driver StabilityWHQL certification for Windows 7 and Windows 8

System Compatibility

Required PCIePCI Express Gen3 (PCI Express 3.0 ×16)
MotherboardLow-profile PCI Express; PCI Express 3.0 ×16
BIOS LimitsUEFI Supported
OS CompatMicrosoft Windows 7; Windows 8; Windows 8.1; Windows 10+; Linux

Benchmarks & Throughput

Inference Benchmarks

22 TOPs of INT8 inference; slashes latency by 15X

Multi-GPU Scalability

Scaling Efficiency

Single GPU60X better energy efficiency than CPUs

Scaling Characteristics

Cross-Node Latencyslashes inference latency by 15X in any hyperscale infrastructure
Parallelismdedicated hardware-accelerated decode engine that works in parallel with the GPU doing inference

Workload Readiness

LLM Inference

22 TOPs (INT8)

Scientific Computing

Single-Precision Performance 5.5 TeraFLOPS

Real-Time Serving

Hardware-decode engine capable of transcoding and inferencing 35 HD video streams in real time

Market Authority

Key Strengths

Limitations

Expert Insight

The P4 represents a strategic leap in AI compute. When comparing cloud providers, consider not just the hourly rate, but also the interconnect bandwidth (InfiniBand/NVLink) and regional availability which can significantly impact total cost of ownership for large-scale training.

Glossary Terms

FP32 TFLOPS
VRAM
TDP
Cores
Information updated daily. Cloud pricing subject to vendor availability.