AI Potluck
Model components / Inference code

SGLang

SGLang (LMSYS / UC Berkeley)

High-performance LLM/VLM serving engine built around RadixAttention for automatic prefix-KV reuse, with structured-generation, speculative decoding, and disaggregated-prefill support. Deployed on 400K+ GPUs worldwide and serving trillions of tokens/day; xAI uses it to serve Grok 3, Microsoft Azure to serve DeepSeek R1 on AMD, and adopters include AMD, NVIDIA, AWS, Oracle Cloud, LinkedIn, and Cursor. Joined the PyTorch ecosystem in March 2025; 28K GitHub stars with 450+ contributors.

v0.5.12.post1, released May 26, 2026 (GitHub). Hosted under LMSYS / PyTorch ecosystem. Confirmed to exist as of June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
governance
LMSYS/PyTorch-ecosystem(community)

Adoption

4 high confidence
4.0

Reported generating trillions of tokens/day in production across 400,000+ GPUs; enterprise users include xAI, NVIDIA, AMD, LinkedIn, Cursor, Oracle/Google/Azure/AWS clouds. Used as rollout backend for frontier-model training (verl, slime, etc.). 28.9k stars secondary. Level 4 reflects very large but operator-concentrated deployment footprint.

Capability

5 high confidence
5.0

RadixAttention, structured generation, multi-step pipelines; benchmarked as top-tier vs vLLM/TensorRT-LLM on H100 in 2026 comparisons; 25x perf improvement reported on NVIDIA GB300 NVL72 (Feb 2026). No formal MLPerf submission confirmed for SGLang itself, so basis is feature_matrix plus third-party benchmarks rather than a standardized suite.

Unchanged since 2026-06-09 (last edited, not re-checked)