Together logo

GPU Cloud Provider · San Francisco, CA

Together

Together AI provides large-scale GPU clusters equipped with NVIDIA's latest Blackwell and Hopper architecture GPUs, offering services primarily aimed at AI training and inference workloads. The clusters are interconnected using NVLink and InfiniBand, and they utilize advanced storage solutions and orchestration through Kubernetes and Slurm to deliver specialized and optimized AI computing resources.

GPUs
14
Founded
Undated

GPU Marketplace

$3.19/hour
$3.45/hour
$3.69/hour
$3.99/hour
$3.99/hour
$4.15/hour
$4.99/hour
$5.49/hour
$5.99/hour
$6.79/hour
$7.79/hour
$7.99/hour
$8.19/hour
$8.99/hour

Estimates onlyrates are collected automatically from public provider pages and may be out of date. Prices vary by region, commitment term, and availability, and typically exclude storage, egress, and tax. Confirm current pricing with the provider before purchasing. Last collected .

Company Profile

FoundedUndated
HeadquartersSan Francisco, CA
Legal EntityTogether AI
FundingAnnouncing our Series C.

Infrastructure

GPU FleetNVIDIA HGX H100, NVIDIA HGX H200, NVIDIA HGX B200, NVIDIA HGX B300, NVIDIA GB200 NVL72, NVIDIA GB300 NVL72
Network FabricInfiniBand, Ethernet
Connectivity14.4 Tbps Infiniband
StorageNVMe SSDs, High-performance converged storage, VAST Data, WEKA AI-native storage systems
Bare MetalSingle-tenant GPU instances with guaranteed performance (no sharing), support for custom models, autoscaling & traffic spike handling.
AvailabilityGA

Compute & Deployment

On-DemandPay as you go GPU capacity on an hourly basis.
Reserved InstancesOn-demand hourly rates and reserved capacity
Bare MetalSingle-tenant GPU instances with: * Guaranteed performance (no sharing) * Support for custom models * Autoscaling & traffic spike handling
VM-BasedCustomize a deployment of VM sandboxes for large development environments.
Serverless GPUServerless Inference

GPU Hardware

Multi-GPU NodesNumber of GPUs in one replica of this instance type.
HGX PlatformNVIDIA HGX H100,NVIDIA HGX H200,NVIDIA HGX B200,NVIDIA HGX B300

Pricing Model

Per HourAll prices are per GPU per hour.
Per Minute$0.05
SubscriptionPay as you go GPU capacity on an hourly basis.
Public Pricinghttps://www.together.ai/pricing
Pay-as-you-goPay as you go GPU capacity on an hourly basis.

Performance & Scaling

Elastic ScalingAutoscaling & traffic spike handling
Auto ScalingAutoscaling & traffic spike handling
Perf IsolationGuaranteed performance (no sharing)
Noisy NeighborGuaranteed performance (no sharing)

Developer Experience

OnboardingStart for free, scale on demand.
FrameworksPyTorch
Model MarketplaceA catalog/listing of many models is presented (examples include MiniMax M3, GLM-5.2, Kimi K3, DeepSeek V4 Flash 0731, Qwen3.8-2.4T-A95B, Muse Glimmer 30B, etc.).
DocumentationOpenAPI specification and API reference describing inference instance types (hardware, pricing, regions, headroom).
API FeaturesCLI, SDK, REST API, Terraform provider

Security & Compliance

Security
Customer logos shown in 'Trusted by' sectionPublic OpenAPI repo license URL

Data Center Locations

Coverage

Compliance Regions

Datacenter Locations

Key Strengths

Transparent, flexible pricing across serverless inference, dedicated endpoints, fine-tuning, and GPU clusters.
Single-tenant GPU instances with guaranteed performance (no sharing), support for custom models, and autoscaling & traffic spike handling.
Reserve dedicated capacity in throughput units (PTUs) for predictable throughput.
High-bandwidth, parallel managed storage colocated with compute.
Fine-tuning for open-source models with standard and specialized pricing.

Known Limitations

No spot instance pricing is listed (only on-demand / reserved pricing shown).
Some hardware rates are not shown and require contacting sales ('[Contact us]' / '[Contact sales]' or '—').
Region headroom may be omitted when unavailable.

Additional Information

Support Options

["Kubernetes Dashboard access","Direct SSH access","Support contact options"]

Community

Public GitHub repository (openapi) indicated by license URL.

Core Proposition

Transparent, flexible pricing across serverless inference, dedicated endpoints, fine-tuning, and GPU clusters. Start for free, scale on demand.

Notable Customers

decagon
cursor
vercept
evertune
cohere
deepmind
lg-ai-research
vfs-global
elevenlabs
arcee
captions
cartesia
sk-telekom
mozilla
hedra
cognition
NOUS
KREA
jasper
salesforce
ai2
lmsys
leonado
snorkel
you
weights-biases
zoho
quora
wp
zoom
nexusflow
upstage
neal-fun
wordware
i.am+
pika
Last updated March 2026. Information subject to change.