AI Potluck
Model components / Inference code

Cerebras Inference

Cerebras Systems

Inference service running on the Cerebras Wafer-Scale Engine (WSE-3), a single chip containing 4 trillion transistors with 44GB of on-chip SRAM, eliminating off-chip memory bottlenecks entirely. Achieves extremely high throughput on large models by keeping the entire model on a single wafer. Available via API. Positioned as the fastest inference for large models where the entire model fits on one wafer. Raised $4.3B+, filed for IPO.

Cerebras Inference: proprietary cloud inference on Wafer-Scale Engine (WSE); OpenAI-compatible API, serves open models (GLM 4.7, gpt-oss-120b, Llama 4 Scout). Closed/managed service. Confirmed live June 2026 (cerebras.ai).

Openness

1 high confidence
1.0
engine
proprietary(WSE wafer-scale stack)
source
closed
access
managed-cloud-API-only
license
Proprietary

Proprietary cloud inference on custom wafer-scale hardware, API-only; closed per recipe (cloud inference engines).

Adoption

3 low confidence
3.0

No developer/user count disclosed on the inference page; named users GSK, Perplexity, AlphaSense, Meta (Llama partnership), DeepLearning.AI. Placed at 3 on reported enterprise traction (named-deployment signal) without a verified usage figure; smaller disclosed footprint than Groq's 3M-developer claim.

Capability

5 high confidence
5.0

Throughput frontier-definer in this category: #1 by output speed on the AA provider leaderboard for gpt-oss-120b (~1,726 t/s, ~2.5x next-fastest SambaNova). Clear C5 (the AA provider-speed leaderboard is the MLPerf analog for inference engines). Live figure drifts run-to-run (~1,726-1,777 t/s); #1 ranking stable.

Unchanged since 2026-06-09 (last edited, not re-checked)