GPU Marketplace

NVIDIA A10On-Demand

NVIDIA L40SOn-Demand

NVIDIA H100 SXMOn-Demand

NVIDIA H200 SXMOn-Demand

NVIDIA B200 SXMOn-Demand

NVIDIA B300 SXMOn-Demand
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Company Profile
Company TypeCloud infrastructure designed for AI workloads
Provider TypeCloud infrastructure designed for AI workloads
Infrastructure
GPU FleetNvidia B300, Nvidia B200, Nvidia H200 SXM, Nvidia H100 SXM5, Nvidia RTX PRO 6000, Nvidia A100, 80 GB, Nvidia A100, 40 GB, Nvidia L40S, Nvidia A10, Nvidia L4, Nvidia T4
Total GPU CapacityScale to 1000+ GPUs in minutes
Built for small teams and independent developersBuilt for startups and larger organizationsFor organizations prioritizing security, support, and everlasting confidence.
Compute & Deployment
Container-BasedAI-native container runtime
Serverless GPUModal is serverless
GPU Hardware
Latest GenNvidia B300, Nvidia B200, Nvidia H200 SXM, Nvidia H100 SXM5, Nvidia RTX PRO 6000, Nvidia A100, 80 GB, Nvidia A100, 40 GB, Nvidia L40S, Nvidia A10, Nvidia L4, Nvidia T4
Pool SizeScale to 1000+ GPUs in minutes.
PCIe vs SXMSXM (listed GPUs: Nvidia H200 SXM; Nvidia H100 SXM5)
Pricing Model
SubscriptionWith Modal, you always pay for what you use and nothing more. You never pay for idle resources — just actual compute time, by the CPU cycle.
Public Pricinghttps://modal.com/pricing
Pay-as-you-goWith Modal, you always pay for what you use and nothing more. You never pay for idle resources — just actual compute time, by the CPU cycle.
Credit SystemGet started with $30 / month free credit
Performance & Scaling
Multi-Node TrainingScale to 1000+ GPUs in minutes. Then back down to zero.
Max Cluster SizeScale to 1000+ GPUs in minutes. Then back down to zero.
Elastic ScalingBurst to thousands of GPUs when demand spikes, then drop back to zero when it doesn’t, keeping workloads efficient.
Auto ScalingModal is serverless, which means that we instantly autoscale up and down for you based on request volume.
Perf IsolationEfficient batching and scheduling keep GPUs near fully loaded, even with bursty or uneven traffic, delivering 2–3× higher throughput per GPU compared to static clusters.
Developer Experience
OnboardingStarter $0 plan with free compute credits; Get started with $30 / month free credit
JupyterModal Sandbox + Notebooks
DocumentationFrequently asked questions
Security & Compliance
SOC 2 complianceHIPAA compatibilityAudit logsSupport via private SlackAWS and GCP marketplace support
Data Center Locations
Coverage
Compliance Regions
Datacenter Locations
Key Strengths
Memory snapshotting for fast model loads
smarter filesystem with lazy loading
serverless autoscaling to thousands of GPUs
deep multi-cloud GPU capacity pool
near-max GPU utilization
per-CPU-cycle billing
Additional Information
Community
Modal Community Slack; Runtime conference for engineers
Core Proposition
You never pay for idle resources — just actual compute time, by the CPU cycle.
Last updated August 2026. Information subject to change.


