Cerebras Inference
Cerebras SystemsCerebras Inference runs on the Cerebras Wafer-Scale Engine (WSE-3), a single silicon wafer with 4 trillion transistors and 44GB of on-chip SRAM. It is delivered as a hosted, API-only cloud service that serves a catalog of models through an OpenAI-compatible API. Cerebras Systems builds and operates the underlying hardware and service.
Verified 2026-08-09 via the Cerebras inference page, the Cerebras chip page, and the Cerebras Inference Terms and Conditions.
Openness
1 high confidence- engine
- proprietary(WSE wafer-scale stack)
- source
- closed
- access
- managed-cloud-API-only
- license
- Proprietary
Proprietary cloud inference on custom wafer-scale hardware, reachable only through an API. The inference page offers managed API access with no source or weights release, and the Inference Terms and Conditions grant nothing more than a non-exclusive, limited, non-transferable and non-sublicensable right to access the hosted services.
- https://www.cerebras.ai/inference recorded 2026-08-09
Cerebras Inference is a hosted cloud service (tiers: Free Trial with "Get API Key", Developer pay-per-token, Enterprise), with "OpenAI API compatibility" advertised as the integration path; a full-text read of the page finds no source repository, no downloadable weights and no self-hostable implementation - the Enterprise tier's "custom model weights" refers to customer-supplied weights served on Cerebras hardware, not a weights release.
- https://d7umqicpi7263.cloudfront.net/eula/4rCqa7snChhObjYFAYH325D2kcqkCRpGXBZYsSMc3mg recorded 2026-08-09
we hereby grant to you or your Named Users a non-exclusive, limited, non-transferable, non-sublicensable right to access and use the Services, solely for your or their internal business purposes
Adoption
3 low confidenceNo developer or user count is disclosed on the inference page. GSK, Perplexity, AlphaSense and DeepLearning.AI are featured as named customer testimonials, and level 3 rests on that reported enterprise traction rather than on a verified usage figure, a smaller disclosed footprint than Groq's published developer count.
- https://www.cerebras.ai/inference recorded 2026-08-09
customer testimonials from GSK ("With Cerebras' inference speed, GSK is developing innovative AI applications"), Perplexity, and AlphaSense (Raj Neervannan, CTO and co-founder); no aggregate developer or user count published
Capability
5 high confidenceDefines the throughput frontier in this category: fastest provider on Artificial Analysis's gpt-oss-120b provider leaderboard at roughly 1,942 output tokens per second, about 2.75x the next-fastest provider, SambaNova, at roughly 705.7, with Groq third at roughly 476.8. Cerebras also claims up to 15x faster than NVIDIA GPUs. That leaderboard is the closest analog to MLPerf for hosted inference engines, and the size of the gap is itself the evidence. The live figure drifts from run to run, an earlier reading being about 1,726 tokens per second, but the first-place ranking has been stable.
- https://artificialanalysis.ai/models/gpt-oss-120b/providers recorded 2026-08-09
Fastest # 1 Cerebras 1,942.0 t/s # 2 SambaNova 705.7 t/s # 3 Groq 476.8 t/s # 4 Google Vertex 395.8 t/s # 5 Azure 323.2 t/s
- https://www.cerebras.ai/inference recorded 2026-08-09
Get instant access to inference that's up to 15x faster than NVIDIA GPUs
Verified 2026-08-09