Together Fine-Tuning
Together AIManaged fine-tuning across 50+ open-weight models (Llama, Mistral, Qwen, DeepSeek) supporting LoRA, full-parameter SFT, tool-call/reasoning-trace/VLM tuning, and models up to 1T params; resulting checkpoints served at base-model inference rates. Differentiates from Fireworks with full fine-tuning (not just LoRA) and broader catalog. Together AI raised $305M Series B at $3.3B valuation (Feb 2025), hit $1B ARR by Feb 2026, with Salesforce, Zoom, and The Washington Post among 450K+ developer accounts.
Together AI managed fine-tuning (docs live June 2026). Methods: LoRA, full fine-tuning, preference (DPO), function-calling, reasoning, and vision-language FT. Can fine-tune 100B+ open models incl. DeepSeek-V3 and Qwen3-235B; handles data prep -> training -> dedicated-endpoint hosting. Proprietary managed service over open base models.
Openness
1 high confidence- source
- closed
- training-pipeline
- closed
- runs-on
- Together-cloud-only
- license
- Proprietary(managed service, per-token training pricing)
- base weights
- open-models(varies)
The fine-tuning service/pipeline is proprietary and runs only on Together's cloud; no source. (It tunes open-weight base models, but the offering itself is closed.)
- https://docs.together.ai/docs/fine-tuning-overview recorded 2026-06-04
managed FT service (LoRA/full/preference/vision); data-prep to hosting; proprietary cloud
Adoption
not assessedNo disclosed standalone usage figure for the Together fine-tuning feature located this run (jobs run / customers tuning). Together AI raised a $305M Series B and serves substantial inference traffic, and a vendor case study cites a customer moving to daily iteration / 77%->87% accuracy, but no hard per-SKU FT usage number was found; declining to assign a level rather than rely on a single anecdote.
- https://docs.together.ai/docs/fine-tuning-overview recorded 2026-06-04
managed FT feature; no aggregate usage figures disclosed
- https://www.together.ai/fine-tuning recorded 2026-06-04
FT product page; case-study traction (daily iteration, 77%->87% accuracy)
Capability
4 medium confidenceBroad managed method coverage (LoRA + full FT + preference + multimodal) over large open models incl. demonstrated 100B+/235B-scale tuning - the widest open-model scale among the managed services here. Black-box (limited scale/precision control), no MLPerf-Training basis, so capped below OSS scale definers; scored 4, on par with the OpenAI/Azure FT APIs.
- https://docs.together.ai/docs/fine-tuning-overview recorded 2026-06-04
LoRA, full FT, preference, function-calling, reasoning, vision-language FT
- https://www.together.ai/fine-tuning recorded 2026-06-04
fine-tune 100B+ models incl. DeepSeek-V3 and Qwen3-235B
Unchanged since 2026-06-09 (last edited, not re-checked)