AI Potluck
Model components / Inference code

Together Inference Engine

Together AI

Proprietary inference engine optimized for serving open source models at scale with Flash Attention, speculative decoding, and custom CUDA kernels. Supports a wide catalog of open models with fast switching. Also offers a dedicated endpoints product for reserved GPU capacity. Founded by researchers from Stanford, Berkeley, and CMU. Differentiated by combining inference speed with strong research output (RedPajama, Together MoE research).

Together AI Inference Engine: proprietary cloud inference (serverless/dedicated) for open models, OpenAI-compatible; built on Together's research stack (FlashAttention, ThunderKittens, ATLAS accelerators). Closed/managed service. Confirmed live June 2026 (together.ai).

Openness

1 high confidence
1.0
engine
proprietary(Together Inference Engine)
source
closed(though Together publishes OSS research e.g. FlashAttention/ThunderKittens separately)
access
managed-cloud-API-only
license
Proprietary

The inference engine/service itself is proprietary, API-only; closed per recipe. (Together open sources research components like FlashAttention separately, but those are not the serving engine.)

  • https://www.together.ai/ recorded 2026-06-04

    proprietary serverless/dedicated cloud inference for open models; engine not released; OSS research published separately

Adoption

4 medium confidence
4.0

Broad named-customer roster (Cursor, Cohere, DeepMind, ElevenLabs, Salesforce, Mozilla, Decagon, ~30+ orgs); widely used serverless inference for open models. No raw developer/user count disclosed; reported_traction at the 1-10M-equivalent band on enterprise breadth.

  • https://www.together.ai/ recorded 2026-06-04

    named customers Cursor, Cohere, DeepMind, ElevenLabs, Salesforce, Mozilla, Decagon + ~30 orgs

Capability

4 low confidence
4.0

Strong GPU-based inference with leading-research kernels, but no Together figure available on the AA gpt-oss-120b provider speed board this run (only TTFT ~4.19s shown, not output t/s), so capability rests on vendor-reported research speedups -> feature_matrix, low confidence. C4: strong GPU provider, not a demonstrated throughput frontier-definer.

Unchanged since 2026-06-24 (last edited, not re-checked)