GPU Marketplace

NVIDIA A30On-Demand
$0.39/hour

NVIDIA A16On-Demand
$0.56/hour

NVIDIA RTX A6000On-Demand
$0.63/hour

NVIDIA GeForce RTX 4090On-Demand
$0.66/hour

NVIDIA GeForce RTX 5090On-Demand
$0.72/hour

NVIDIA RTX 4000 AdaOn-Demand
$0.87/hour

NVIDIA RTX A4000On-Demand
$0.88/hour

NVIDIA L40SOn-Demand
$0.97/hour

NVIDIA L40On-Demand
$0.97/hour

NVIDIA RTX 6000 AdaOn-Demand
$1.07/hour

NVIDIA A10On-Demand
$1.42/hour

NVIDIA RTX A5000On-Demand
$1.56/hour

NVIDIA A40On-Demand
$2.05/hour
$2.41/hour
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Company Profile
Provider TypeGPU compute
Legal EntityRuncrate
Infrastructure
GPU Fleet22 GPU SKUs available
Bare MetalInstances run as containers, VMs, or bare-metal
Self-ServeDedicated, and Enterprise
Compute & Deployment
On-DemandGPU compute on demand
Reserved InstancesReserved GPU capacity, no scale-to-zero
Bare MetalInstances run as containers, VMs, or bare-metal
VM-BasedInstances run as containers, VMs, or bare-metal
Container-BasedInstances run as containers, VMs, or bare-metal
Spin-Up Time<60s
GPU Hardware
Latest GenB200, H200, H100, A100, L40S, RTX 6000 Ada
Multi-GPU Nodessingle-node to 8× H200 clusters
Pricing Model
SubscriptionSelf-Serve; Dedicated; Enterprise
Reserved Discount40–60% off the public rate card
Public PricingPublic rate card on every open-source model.
Pay-as-you-gopay as you go
Performance & Scaling
Multi-Node Trainingsingle-node to 8× H200 clusters
Elastic ScalingMulti-cloud autoscaling
Auto Scalingmulti-cloud autoscaler
Developer Experience
OnboardingSelf-Serve: Get an API key, ship today, no sales contact; Live API endpoint available with no key required; Dedicated: 7-day pilot parallel to current provider
JupyterJupyterLab available as a template
TemplatesPyTorch, vLLM, Axolotl, ComfyUI, JupyterLab
Model MarketplacePublic rate card on every open-source model; 170+ models listed
DocumentationReal-time GPU pricing API endpoint (/api/gpu-pricing); FAQ with common questions; Templates library with PyTorch, vLLM, Axolotl, ComfyUI, JupyterLab
Security & Compliance
SOC 2 Type II datacenter partnersHIPAA-eligiblenamed model partners listed (e.g.MetaOpenAIAlibaba CloudNVIDIA)
Data Center Locations
Coverage
CountriesUS
Latency Tiersp99 latency floor
Multi-cloud · 12 regions · 24/7 capacity
Compliance Regions
EU Data ResidencyRegion pinning — US / EU / APAC
Datacenter Locations
Key Strengths
Per-token inference on open-source models with a public rate card
per-second GPU compute and per-second billing
multi-cloud across 12 regions with <60s cold starts via template cache
Instances and Crates (containers
VMs, bare-metal)
broad GPU lineup (L40S to B200
22 SKUs)
Dedicated reserved capacity with 40–60% discounts and contractual p99/p99.9+ SLAs
BYOC/self-hosted and SOC 2 Type II datacenter partners
Additional Information
Community
Email + community Discord; Slack Connect option
Core Proposition
Pay-per-token inference on every open-source model. Per-second GPU compute.
Last updated March 2026. Information subject to change.

