Together Inference Engine
Together AIProprietary inference engine optimized for serving open source models at scale with Flash Attention, speculative decoding, and custom CUDA kernels. Supports a wide catalog of open models with fast switching. Also offers a dedicated endpoints product for reserved GPU capacity. Founded by researchers from Stanford, Berkeley, and CMU. Differentiated by combining inference speed with strong research output (RedPajama, Together MoE research).
Together AI Inference Engine: proprietary cloud inference (serverless/dedicated) for open models, OpenAI-compatible; built on Together's research stack (FlashAttention, ThunderKittens, ATLAS accelerators). Closed/managed service. Confirmed live June 2026 (together.ai).
Openness
1 high confidence- engine
- proprietary(Together Inference Engine)
- source
- closed(though Together publishes OSS research e.g. FlashAttention/ThunderKittens separately)
- access
- managed-cloud-API-only
- license
- Proprietary
The inference engine/service itself is proprietary, API-only; closed per recipe. (Together open sources research components like FlashAttention separately, but those are not the serving engine.)
- https://www.together.ai/ recorded 2026-06-04
proprietary serverless/dedicated cloud inference for open models; engine not released; OSS research published separately
Adoption
4 medium confidenceBroad named-customer roster (Cursor, Cohere, DeepMind, ElevenLabs, Salesforce, Mozilla, Decagon, ~30+ orgs); widely used serverless inference for open models. No raw developer/user count disclosed; reported_traction at the 1-10M-equivalent band on enterprise breadth.
- https://www.together.ai/ recorded 2026-06-04
named customers Cursor, Cohere, DeepMind, ElevenLabs, Salesforce, Mozilla, Decagon + ~30 orgs
Capability
4 low confidenceStrong GPU-based inference with leading-research kernels, but no Together figure available on the AA gpt-oss-120b provider speed board this run (only TTFT ~4.19s shown, not output t/s), so capability rests on vendor-reported research speedups -> feature_matrix, low confidence. C4: strong GPU provider, not a demonstrated throughput frontier-definer.
- https://www.together.ai/ recorded 2026-06-04
FlashAttention-4 1.3x cuDNN on Blackwell, ATLAS up to 4x LLM inference, 2x-faster claim
- https://artificialanalysis.ai/models/gpt-oss-120b/providers recorded 2026-06-04
Together.ai present on provider board (TTFT ~4.19s); output t/s not in displayed ranking
Unchanged since 2026-06-24 (last edited, not re-checked)