Cerebras Inference
Cerebras SystemsInference service running on the Cerebras Wafer-Scale Engine (WSE-3), a single chip containing 4 trillion transistors with 44GB of on-chip SRAM, eliminating off-chip memory bottlenecks entirely. Achieves extremely high throughput on large models by keeping the entire model on a single wafer. Available via API. Positioned as the fastest inference for large models where the entire model fits on one wafer. Raised $4.3B+, filed for IPO.
Cerebras Inference: proprietary cloud inference on Wafer-Scale Engine (WSE); OpenAI-compatible API, serves open models (GLM 4.7, gpt-oss-120b, Llama 4 Scout). Closed/managed service. Confirmed live June 2026 (cerebras.ai).
Openness
1 high confidence- engine
- proprietary(WSE wafer-scale stack)
- source
- closed
- access
- managed-cloud-API-only
- license
- Proprietary
Proprietary cloud inference on custom wafer-scale hardware, API-only; closed per recipe (cloud inference engines).
- https://www.cerebras.ai/inference recorded 2026-06-04
proprietary cloud inference on Wafer-Scale Engine, OpenAI-compatible API, serves GLM/gpt-oss/Llama Scout
Adoption
3 low confidenceNo developer/user count disclosed on the inference page; named users GSK, Perplexity, AlphaSense, Meta (Llama partnership), DeepLearning.AI. Placed at 3 on reported enterprise traction (named-deployment signal) without a verified usage figure; smaller disclosed footprint than Groq's 3M-developer claim.
- https://www.cerebras.ai/inference recorded 2026-06-04
named users GSK, Perplexity, AlphaSense, Meta, DeepLearning.AI; no user-count figure
Capability
5 high confidenceThroughput frontier-definer in this category: #1 by output speed on the AA provider leaderboard for gpt-oss-120b (~1,726 t/s, ~2.5x next-fastest SambaNova). Clear C5 (the AA provider-speed leaderboard is the MLPerf analog for inference engines). Live figure drifts run-to-run (~1,726-1,777 t/s); #1 ranking stable.
- https://artificialanalysis.ai/models/gpt-oss-120b/providers recorded 2026-06-04
Cerebras #1 output speed ~1,726 t/s for gpt-oss-120b; next-fastest SambaNova ~692.5 t/s (~2.5x)
- https://www.cerebras.ai/inference recorded 2026-06-04
up to 15x faster than NVIDIA GPUs; >2,000 t/s on Llama 4 Scout
Unchanged since 2026-06-09 (last edited, not re-checked)