Together AI
Together AI is a cloud platform providing infrastructure for running, fine-tuning and training AI models, including serverless inference APIs, dedicated GPU instances and full GPU clusters. Serverless inference is priced per token, from $0.09 to $3.00 per million input tokens depending on the model, with dedicated GPU instances from $3.99 an hour.
Facts checked on against the sources listed at the end of this page.
- Pricing
- Pay as you goPaid from $0.09 per million input tokens (lowest-cost serverless models)
- Open source
- NoProprietary (platform); hosts many open source and open-weight models
- npm downloads a week
- 120kup 8% on the week before
A good fit for
- Developers wanting API access to open source and open-weight models without self-hosting them
- Teams needing GPU infrastructure for training or fine-tuning custom models
- Cost-sensitive projects that can take advantage of cached-token discounts and off-peak or batch pricing
- Applications needing dedicated, SLA-backed inference capacity rather than shared serverless endpoints
Not a good fit for
- Teams wanting a single proprietary flagship model rather than a broad catalog of hosted open models
- Very small projects that don't need dedicated GPU capacity or fine-tuning infrastructure
- Users wanting fully predictable flat pricing rather than usage-based, model-dependent rates
Together AI pricing
Serverless chat model inference ranges from $0.09 to $3.00 per million input tokens depending on the model, with output tokens typically costing 3 to 5 times more than input tokens; cached tokens get up to 90% discounts. Dedicated GPU instances range from $3.99 to $9.99 per GPU per hour on demand, with preemptible GPU cluster compute starting at $1.99 an hour and reserved options for 181+ days. Provisioned throughput starts around $21,600 a month for baseline capacity. Fine-tuning costs range from $0.34 to $40 per million tokens depending on model size and training type.
| Plan | Price | What you get |
|---|---|---|
| Serverless inference | $0.09-$3.00/M input tokens (model-dependent) | Pay-per-token API access to a wide range of hosted models. |
| Dedicated inference | $3.99-$9.99/GPU/hour (on-demand) | Reserved GPU instances for dedicated model serving. |
| GPU clusters | From $1.99/hour (preemptible) | Rentable GPU hardware for custom training and inference workloads. |
Prices in USD, checked on 27 September 2026. Current prices on together.ai.
Main features
- Serverless inference
- Pay-per-token API access to a wide range of open source and open-weight models.
- Dedicated inference
- Reserved GPU instances for consistent, dedicated model serving.
- GPU clusters
- Rent GPU hardware (H100 through GB300) for custom training or inference workloads.
- Fine-tuning
- LoRA and full fine-tuning services for adapting models to specific use cases.
- Batch inference
- Cost-efficient processing for large-scale, non-real-time inference workloads.
About Together AI
Together AI runs and hosts a wide range of open source and open-weight AI models, giving developers API access to models they'd otherwise have to self-host, alongside infrastructure for training and fine-tuning their own models.
Its offerings span serverless inference (pay-per-token API access), provisioned throughput with SLA guarantees, dedicated inference on reserved GPU hardware, and raw GPU cluster rental for teams that want to run their own training or inference workloads directly.
Pricing reflects this range of services: serverless chat model inference is priced per million tokens and varies by model, while GPU cluster and dedicated inference pricing is billed by the hour or through longer-term reservations.
- Made by
- Together AI
- Free plan
- Limited free credits are typically available for new accounts to try the platform
- Available on
- API, GPU cloud (H100, H200, B200, B300, GB200, GB300), Fine-tuning and training infrastructure
Usage and activity
Weekly npm downloads of together-ai, last 12 months.
Together AI alternatives
Questions about Together AI
What is Together AI?
A cloud platform for running, fine-tuning and training AI models, offering serverless inference, dedicated GPU instances and GPU clusters.
How is Together AI priced?
Serverless inference is priced per million tokens and varies by model, from $0.09 to $3.00 per million input tokens, while GPU and dedicated inference services are billed hourly or through reservations.
Does Together AI host open source models?
Yes, it hosts a wide range of open source and open-weight models for API access, in addition to offering infrastructure for custom training.
Can I fine-tune models on Together AI?
Yes, it offers LoRA and full fine-tuning services, priced from $0.34 to $40 per million tokens depending on model size and training type.
Sources and links
Work on Together AI? Claim this listing. Something wrong or out of date? Suggest a correction. Stats come from the public GitHub and npm APIs and update every day.
