Replicate logo

GPU Cloud Provider

Replicate

GPUs
2

GPU Marketplace

NVIDIA T4On-Demand
$0.81/hour
NVIDIA L40SOn-Demand
$3.51/hour

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Company Profile

Legal EntityReplicate

Infrastructure

GPU FleetProvides instances using Nvidia A100 (80GB), H100, L40S, T4, and H200 GPUs (single- and multi-GPU instance types listed)
Bare MetalMost private models run on dedicated hardware
Enterprise & volume discounts

Compute & Deployment

Reserved Instancescommitted spend contracts
Bare Metalmost private models run on dedicated hardware
Serverless GPUfast booting fine-tunes billed only when active

GPU Hardware

Multi-GPU NodesYes

Pricing Model

SubscriptionYou only pay for what you use on Replicate. Some models are billed by time, others by input and output.
Public Pricinghttps://replicate.com/pricing
Pay-as-you-goYou only pay for what you use on Replicate.

Performance & Scaling

Elastic ScalingIf you get a ton of traffic, we automatically scale up and down to handle the demand.
Auto ScalingIf you get a ton of traffic, we automatically scale up and down to handle the demand.
Perf IsolationUnlike public models, most private models (with the exception of fast booting fine-tunes(https://replicate.com/docs/billing#fast-booting-fine-tunes)) run on dedicated hardware so you don't have to share a queue with anyone else.
Noisy NeighborUnlike public models, most private models (with the exception of fast booting fine-tunes(https://replicate.com/docs/billing#fast-booting-fine-tunes)) run on dedicated hardware so you don't have to share a queue with anyone else.

Developer Experience

OnboardingEnterprise offers help with onboarding
SDK LanguagesNode.js and Python (documentation pages exist)
JupyterGoogle Colab support documented
Model MarketplacePublic models: thousands of community-contributed open-source models and proprietary models hosted
DocumentationDocs include getting-started guides, guides, topics, and a reference with client libraries and HTTP API documentation

Security & Compliance

Hosts models from named providers (examples shown on pricing page include anthropic, black-forest-labs, ideogram-ai, deepseek-ai, wavespeedai, recraft-ai)

Data Center Locations

Coverage

Compliance Regions

Datacenter Locations

Key Strengths

Pay-per-use billing
large public model catalog plus private model deployments
private models on dedicated hardware
fast-booting fine-tunes billed only while active
wide range of GPU instance types (A100/H100/L40S/T4/H200)

Known Limitations

No API currently to change hardware for public/private models
private models (non-fast-booting) are billed for idle time
additional multi-GPU / H200 capacity requires committed spend contracts

Additional Information

Community

Open-source tooling (Cog) and documentation; public docs and references

Core Proposition

You only pay for what you use on Replicate. Some models are billed by hardware and time, others by input and output.

Last updated August 2026. Information subject to change.