AI Potluck
Model components / Inference code

Groq Inference

Groq

Inference service running on custom Language Processing Unit (LPU) hardware designed from the ground up for sequential token generation. Achieves industry-leading token throughput, widely benchmarked at 500-800 tokens/sec for Llama-class models, far exceeding GPU-based alternatives. Deterministic, SRAM-based architecture eliminates memory bottlenecks. Available via API with OpenAI-compatible endpoints. The company that put 'inference speed' in the mainstream conversation in early 2024.

GroqCloud proprietary cloud inference on Groq LPU (Language Processing Unit) hardware; OpenAI-compatible API, serves open models (gpt-oss, Llama, etc.). Closed/managed-only service. Confirmed live June 2026 (groq.com).

Openness

1 high confidence
1.0
engine
proprietary(LPU stack)
source
closed
access
managed-cloud-API-only
license
Proprietary

Proprietary cloud inference engine + custom LPU silicon, API-only; per recipe, cloud inference engines (Groq) are closed.

  • https://groq.com/ recorded 2026-06-04

    GroqCloud proprietary managed inference on LPU hardware, API access via console, OpenAI-compatible

Adoption

4 medium confidence
4.0

Vendor reports '3m developers and teams' on GroqCloud; named enterprise users incl. Dropbox, Vercel, Canva, Robinhood, Volkswagen, Workday, Ramp. Developer-count is vendor-self-reported (reported_traction), placing it in the 1-10M band.

  • https://groq.com/ recorded 2026-06-04

    '3m developers and teams' use Groq; named customers Dropbox/Vercel/Canva/Robinhood/Workday/Ramp

Capability

5 medium confidence
5.0

Capability for an inference engine = throughput/latency frontier (the MLPerf analog here is the Artificial Analysis provider-speed leaderboard). Groq LPU is a recognized speed-frontier provider class; C5 alongside the open-engine anchors. NOTE: could not pull Groq's exact t/s figure for gpt-oss-120b on the AA page this run (Groq present but not in the displayed top-5); score leans on its established speed-leadership positioning -> medium confidence.

Unchanged since 2026-06-09 (last edited, not re-checked)