Wednesday, May 20, 2026
S&P 500 · NDX · NVDA · Gold · BTC
  Tool · Compute & Inference All tools →
Together AI

Together AI.

Inference and fine-tuning compute for open models — the open-source AWS for AI.

Category
Compute & Inference
Vendor
Together AI
Domain
together.ai
Pricing
Pay-per-token inference · Dedicated GPU clusters · Enterprise
Region
San Francisco, US
Models
Llama, Mistral, Qwen, DeepSeek, Mixtral — 200+ open models served

What it is.

San Francisco compute platform built to serve open-weights models at frontier-closed quality and a fraction of the price. Founded by the founder and a roster of distributed-systems researchers. Raised over $300M from Salesforce Ventures, Kleiner Perkins, and NVIDIA. Operates one of the largest open-model inference fleets outside the hyperscalers.

Where it fits.

Engineering teams running production traffic on Llama, Mistral, or Qwen who need predictable per-token pricing and dedicated capacity. Default landing spot when a buyer's Bedrock or Vertex bill compounds past $50,000 a month. Strong fit for fine-tuning workflows where the team wants to keep the resulting weights.

Use cases
  • Open-model inference at production scale
  • Fine-tuning Llama and Mistral on proprietary data
  • Dedicated GPU cluster rental
  • Cost-controlled inference for high-volume workloads
Strengths
  • Per-token cost typically 30 to 60 per cent below Bedrock and Vertex
  • Dedicated cluster pricing predictable at production scale
  • Custom fine-tunes hosted alongside base models on the same API
Considerations
  • Frontier-closed models from OpenAI and Anthropic not available
  • SOC 2 governance maturity behind hyperscaler equivalents

Frequently asked.

How does Together compare to AWS Bedrock?

Lower per-token cost on open models and faster shipping of new open weights. Bedrock wins on procurement integration and on serving frontier-closed Anthropic models.

Can I fine-tune a model and keep the weights?

Yes. Fine-tunes on Together can be downloaded as standard weights and served anywhere. This is the central differentiator versus Bedrock and Vertex.

What is the typical inference latency?

First-token latency 150 to 300 milliseconds on shared endpoints. Dedicated clusters with co-located TTS run inside 100 milliseconds.