AI Potluck
Model components / Inference code

SambaNova Cloud

SambaNova Systems

Inference service running on custom Reconfigurable Dataflow Architecture (RDA) chips, the SN40L processor with 1TB+ on-chip memory. Designed for large models that don't fit in GPU memory without tensor parallelism. Offers both cloud API and on-premise appliance (DataScale) for enterprise deployments. Competitive with Groq and Cerebras on throughput for large models. Raised $1.5B+, primarily targeting enterprise and government customers.

SambaNova Cloud, proprietary managed inference service on SN40L RDU hardware (cloud.sambanova.ai). Serves Llama / open models at high token/s. Confirmed as live commercial cloud offering June 2026.

Openness

1 high confidence
1.0
license
proprietary
software-stack
closed(SambaFlow compiler/runtime not open)
hardware
SN40L-RDU-only
delivery
managed-cloud-API

Proprietary managed cloud inference on SambaNova's own RDU silicon; no open source software disclosed. Scored closed.

Adoption

3 low confidence
3.0

Marketed as fast inference cloud with free tier; integrated by tools (Vellum) and benchmarked by third parties, but a niche specialized provider vs the hyperscalers. No verified user count; level 3 = reported traction.

Capability

4 medium confidence
4.0

Among the fastest per-request token/s in market on its silicon (full-precision 405B at 129 tok/s, 8B >1000 tok/s). Vendor/3rd-party throughput figures; specialized hardware, narrower model/feature breadth than the open frontier engines, so C4 not C5.

Unchanged since 2026-06-24 (last edited, not re-checked)