AI Potluck
Model components / Inference code

Fireworks Inference

Fireworks AI

Fireworks Inference runs open models on FireAttention, an in-house attention engine of custom CUDA and AMD-GPU kernels, combined with speculative decoding. It hosts many models and LoRA adapters concurrently, supports function calling, and offers both serverless per-token endpoints and dedicated GPU deployments. Its vendor, Fireworks AI, was founded by former Meta PyTorch engineers, including CEO Lin Qiao.

Verified 2026-08-09 via the Fireworks homepage, team page, and FireAttention V3 blog post.

Openness

1 high confidence
1.0
engine
proprietary(FireAttention kernels)
source
closed
access
managed-cloud-API-only
license
Proprietary

A proprietary inference engine built on Fireworks' own FireAttention kernels and offered only as a managed cloud service. The fireworks.ai homepage carries no self-host, on-premises, VPC or BYOC option and no repository link anywhere on the page, and no LICENSE or license grant is published for the engine, so there is no source to run yourself and nothing to license.

  • https://fireworks.ai/ recorded 2026-08-09

    the site markets Fireworks purely as a hosted API/pricing product (serverless and on-demand tiers, /pricing links); no self-host, on-prem, VPC, BYOC, or GitHub/source-repo link appears anywhere on the page

Adoption

4 medium confidence
4.0

A strong named-customer roster. Cursor, Vercel, Notion, Sourcegraph, UiPath, Quora, Cresta and StackBlitz all appear on the homepage with logos and case studies, and Fireworks serves top open models behind a number of high-traffic AI products. The Notion case study quotes AI lead Sarah Sachs on latency falling from about 2 seconds to 350 milliseconds, and the Quora one quotes product lead Spencer Chan on a 3x speedup in response time. No raw user count is disclosed, so the band rests on the breadth of high-volume production customers, which supports the 1M-10M equivalent range.

  • https://fireworks.ai/ recorded 2026-08-09

    customer logos and quotes for Cursor, Vercel, Notion, Sourcegraph, UiPath, Quora, Cresta, StackBlitz; Notion (Sarah Sachs, AI Lead): "we reduced latency from about 2 seconds to 350 milliseconds"; Quora (Spencer Chan, Product Lead): "we noticed a 3x speedup in response time"

Capability

4 medium confidence
4.0

Top-tier GPU-based inference on Fireworks' FireAttention kernels, sitting just below the wafer-scale and LPU speed frontier. The band rests on Artificial Analysis's Fireworks provider page, which carries current figures for the ten models Fireworks serves and tops out at 459 output tokens per second. Fireworks does not appear on Artificial Analysis's per-model gpt-oss-120b providers board, which says something about that board's coverage rather than about the product; that board is still cited because it is what establishes where the speed frontier sits, with Cerebras at 1,820.3, SambaNova at 701.3 and Groq at 475.9 output tokens per second.

  • https://artificialanalysis.ai/providers/fireworks recorded 2026-08-13

    10 models served; output speed 459 t/s (Nemotron 3.5 Lightning BF16), 206 t/s (MiniMax-M2.7), 115 t/s (DeepSeek V4 Flash), 107 t/s (Kimi K2.6); latency 0.93s best (Kimi K2.6); "448% difference between the fastest and slowest"

  • https://artificialanalysis.ai/models/gpt-oss-120b/providers recorded 2026-08-13

    The speed frontier this band is placed against: Cerebras 1,820.3, SambaNova 701.3, Groq 475.9, Google Vertex 362.7, Azure 310.5 output t/s - and no Fireworks entry, which is why the provider page rather than this one carries the basis.

  • https://fireworks.ai/ recorded 2026-08-09

    Notion (Sarah Sachs, AI Lead): "we reduced latency from about 2 seconds to 350 milliseconds"; Quora (Spencer Chan, Product Lead): "we noticed a 3x speedup in response time"

Verified 2026-08-09