LMDeploy
Shanghai AI LaboratoryLMDeploy v0.13.0 (released May 12, 2026, GitHub); Apache-2.0 OSS inference/serving toolkit (TurboMind + PyTorch engines). Confirmed live on GitHub + PyPI June 2026.
LMDeploy v0.13.0 (released May 12, 2026, GitHub); Apache-2.0 OSS inference/serving toolkit (TurboMind + PyTorch engines). Confirmed live on GitHub + PyPI June 2026.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(InternLM/lmdeploy)
- quantization
- weight-only+kv(AWQ/W4A16)
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), source public, no proprietary core; broad model coverage (Llama/InternLM/Qwen/DeepSeek/Mixtral + VLMs).
- https://github.com/InternLM/lmdeploy recorded 2026-06-04
Apache-2.0 license, v0.13.0 (May 12 2026), weight-only+kv quantization, Llama/InternLM/Qwen/DeepSeek/Mixtral + VLM support
Adoption
3 medium confidence~165K PyPI downloads last month (pypistats); 352 GitHub dependents. Real usage volume in the 100K-1M band. Smaller footprint than vLLM/SGLang anchors (millions/mo) but well beyond hobbyist; especially used in the InternLM/Chinese-model ecosystem. 7.9k stars corroborate only.
- https://pypistats.org/packages/lmdeploy recorded 2026-06-04
~165,001 downloads last month
- https://github.com/InternLM/lmdeploy recorded 2026-06-04
352 dependent projects, 7.9k stars
Capability
4 medium confidenceStrong, full-featured GPU inference engine (continuous batching, quantization, TP, VLM support). No standardized MLPerf submission confirmed, so feature_matrix not benchmark; a notch below the vLLM/SGLang/TensorRT-LLM C5 anchors on throughput-leadership/ecosystem breadth.
- https://github.com/InternLM/lmdeploy recorded 2026-06-04
TurboMind/PyTorch engines, continuous batching, AWQ/kv quantization, TP, VLM support, 2.4x 4-bit claim
Unchanged since 2026-07-30 (last edited, not re-checked)