GPU Marketplace

NVIDIA A10GOn-Demand

NVIDIA B200 SXMOn-Demand
Estimates only — rates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .
Company Profile
Provider TypeCloud Pricing
Legal EntityBaseten
Infrastructure
GPU FleetNVIDIA T4, NVIDIA L4, NVIDIA A10G, NVIDIA A100, NVIDIA H100, H100 MIG (Fractional NVIDIA H100), NVIDIA H200, NVIDIA B200, NVIDIA RTX-PRO-6000
Bare MetalSelf-host deployments
Production inference that won't break your product or your bank.
Compute & Deployment
On-DemandOn-demand compute
VM-BasedT4 16 GiB VM
GPU Hardware
Multi-GPU NodesFor models that don't fit on a single node, set `node_count` in `resources` to provision multiple identical nodes for one deployment. Each node gets the resources you specify, and Baseten connects the nodes with high-speed InfiniBand for inter-node communication.
InfiniBandBaseten connects the nodes with high-speed InfiniBand for inter-node communication.
Pricing Model
Per Minute$0.00058/min
SubscriptionBasic; Pro; Enterprise
Public Pricinghttps://www.baseten.co/pricing/
Pay-as-you-go$0 per month, pay as you go
Performance & Scaling
Elastic ScalingUnlimited autoscaling and priority compute access
Auto ScalingBest-in-class model performance, effortless autoscaling, and blazing fast cold starts mean you get the most out of each GPU, saving money along the way.
InfiniBandEach node gets the resources you specify, and Baseten connects the nodes with high-speed InfiniBand for inter-node communication.
Developer Experience
OnboardingBasic: $0 per month, pay as you go
CLI ToolingTruss CLI (commands referenced: truss push, truss push --watch, truss push --promote, truss watch, truss push --environment)
Model MarketplaceInstant access to pre-optimized models running on the Baseten Inference Stack.
DocumentationDocumentation Index; instance type reference; config.yaml examples and code snippets; Truss CLI commands and examples; GPU instance tables and GPU details/workloads; multi-node deployment guidance
Security & Compliance
SOC 2 Type II and HIPAA compliant
Data Center Locations
Coverage
Compliance Regions
Datacenter Locations
Key Strengths
Dedicated deployments
Model APIs
Training
Fast cold starts
SOC 2 Type II and HIPAA compliance
Unlimited autoscaling and priority compute access
Priority access to high-demand GPUs
Dedicated compute and higher Model API rate limits
Hands-on engineering expertise and dedicated support on Slack/Zoom
Custom SLAs and self-host deployments
On-demand flex compute and ability to use existing cloud commitments
Full control over data residency, advanced security/compliance, custom global regions, and advanced RBAC
Minute-level billing and multi-node deployments with high-speed InfiniBand.
Known Limitations
Compute in other countries/regions requires contacting sales
L4 GPU is 'Not suitable for LLMs due to bandwidth'
Changes to config.yaml only affect new deployments — updates to existing published deployments must be done through the Baseten UI.
Additional Information
Community
Dedicated support on Slack and Zoom
Core Proposition
Best-in-class model performance, effortless autoscaling, and blazing fast cold starts mean you get the most out of each GPU, saving money along the way.
Last updated August 2026. Information subject to change.

