Together Inference Engine
Together AITogether Inference Engine is a serving stack that runs open models with FlashAttention, adaptive speculative decoding (ATLAS), and custom CUDA kernels. It offers serverless endpoints across a large catalog of open models alongside dedicated instances that reserve GPU capacity for a single model. Together AI builds it, and its team also publishes open kernel research including FlashAttention and ThunderKittens.
Verified 2026-08-09 via together.ai's Serverless Inference and Dedicated Model Inference pages, the FlashAttention-4 and ATLAS blog posts, and the Terms of Service.
Openness
1 high confidence- engine
- proprietary(Together Inference Engine)
- source
- closed(though Together publishes OSS research e.g. FlashAttention/ThunderKittens separately)
- access
- managed-cloud-API-only
- license
- Proprietary
No self-hostable implementation of the Together Inference Engine exists: the serverless and dedicated product pages describe a managed cloud API and "reserved, isolated GPU endpoints running Together's proprietary inference engine". The Terms of Service grant only a non-exclusive, non-sublicensable right to access the service, with Together reserving all intellectual property. Together does publish open-source research such as FlashAttention and ThunderKittens, but that is separate from the engine scored here.
- https://www.together.ai/inference recorded 2026-08-09
"Pay-per-use access to 200+ open-source models across chat, image, video, code, and voice modalities. OpenAI-compatible API, sub-second latency, adaptive speculative decoding, and batch processing at up to 50% lower cost. No infrastructure to manage."
- https://www.together.ai/dedicated-model-inference recorded 2026-08-09
"Reserved, isolated GPU endpoints running Together's proprietary inference engine for production workloads."
- https://www.together.ai/terms-of-service recorded 2026-08-09
"the Company hereby grants you a non-exclusive, non-sublicensable right and license to use the Company IP as permitted by this Agreement. Company reserves all of its intellectual property and other proprietary rights not expressly granted to you herein."
Adoption
4 medium confidenceThe homepage carries a named-customer case study, "How Cursor partnered with Together AI to deliver real-time, low-latency inference at scale", alongside a hero logo collection including Salesforce, ElevenLabs, Cohere, DeepMind, Mozilla and Decagon. No developer or user count is disclosed, so the band rests on that reported enterprise breadth.
- https://www.together.ai/ recorded 2026-08-09
story-card title "How Cursor partnered with Together AI to deliver real-time, low-latency inference at scale"; hero logo images salesforce_black.svg, elevenlabs_black.svg, cohere_black.svg, deepmind_black.svg, mozilla_black.svg, and a Decagon story card
Capability
4 low confidenceThe score rests on vendor-reported speedups taken from Together's own FlashAttention-4 and ATLAS research posts rather than from the homepage summary of them: FlashAttention-4 is reported at 1.1-1.3x faster than cuDNN 9.13 and 2.1-2.7x faster than Triton on B200/Blackwell, up to 1,605 TFLOPs/s at 71% utilization, and ATLAS at up to 4x faster LLM inference. Artificial Analysis's gpt-oss-120b provider board would be the independent check, but its output-speed and time-to-first-token table is rendered client-side and absent from the static page, where "togetherai" appears only in an embedded list of provider names with no figure beside it. With only vendor-reported research figures behind it, confidence is low.
- https://www.together.ai/blog/flashattention-4 recorded 2026-08-09
"FlashAttention-4 is 1.1-1.3x faster than cuDNN 9.13 and 2.1-2.7x faster than Triton"; "it reaches up to 1605 TFLOPs/s (71% utilization)"
- https://www.together.ai/blog/adaptive-learning-speculator-system-atlas recorded 2026-08-09
"ATLAS delivers up to 4x faster LLM inference, powered by Together Turbo's latest research."
Verified 2026-08-09