Together AI.
Inference and fine-tuning compute for open models — the open-source AWS for AI.
What it is.
San Francisco compute platform built to serve open-weights models at frontier-closed quality and a fraction of the price. Founded by the founder and a roster of distributed-systems researchers. Raised over $300M from Salesforce Ventures, Kleiner Perkins, and NVIDIA. Operates one of the largest open-model inference fleets outside the hyperscalers.
Where it fits.
Engineering teams running production traffic on Llama, Mistral, or Qwen who need predictable per-token pricing and dedicated capacity. Default landing spot when a buyer's Bedrock or Vertex bill compounds past $50,000 a month. Strong fit for fine-tuning workflows where the team wants to keep the resulting weights.
- Per-token cost typically 30 to 60 per cent below Bedrock and Vertex
- Dedicated cluster pricing predictable at production scale
- Custom fine-tunes hosted alongside base models on the same API
- Frontier-closed models from OpenAI and Anthropic not available
- SOC 2 governance maturity behind hyperscaler equivalents
Frequently asked.
How does Together compare to AWS Bedrock?
Lower per-token cost on open models and faster shipping of new open weights. Bedrock wins on procurement integration and on serving frontier-closed Anthropic models.
Can I fine-tune a model and keep the weights?
Yes. Fine-tunes on Together can be downloaded as standard weights and served anywhere. This is the central differentiator versus Bedrock and Vertex.
What is the typical inference latency?
First-token latency 150 to 300 milliseconds on shared endpoints. Dedicated clusters with co-located TTS run inside 100 milliseconds.