verl
ByteDance Seed / Volcano EngineHybridFlow flexible and efficient RL post-training framework developed by ByteDance Seed MLSys and HKU, purpose-built for RLHF and reasoning RL on math and code tasks. Picked when teams need state-of-the-art GRPO/PPO at frontier scale; verl shipped early support for DeepSeek-R1-style reasoning RL recipes and has become a reference implementation in the post-DeepSeek wave. 21.5K GitHub stars, 618 contributors, ~420 commits in 90 days, one of the fastest-growing repos in the category.
verl (Volcano Engine RL / HybridFlow), v0.8.0 released June 1 2026. Initiated by ByteDance Seed, now community-maintained. RL post-training library (takes existing weights and adapts via RLHF/RL). Confirmed live on GitHub June 2026.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public
- governance
- community(ByteDance-Seed origin)
- core-gated
- ungated
Fully OSI-licensed; full source public, no managed-tier gating of the library itself.
- https://github.com/volcengine/verl recorded 2026-06-04
Apache-2.0 license; v0.8.0 (Jun 1 2026); RL post-training framework
Adoption
4 high confidence~96.7k PyPI downloads last month (primary). De facto standard high-throughput RL post-training lib; used to train Seed-Thinking-v1.5 and Doubao, and adopted/contributed by Qwen team, Moonshot/Kimi, NVIDIA research, Microsoft Research, Amazon, LinkedIn, Xiaomi, Baidu and many labs. Level 4 reflects broad and growing developer-tool usage at the high end of the 100K-1M band trending toward heavier use.
- https://pypistats.org/packages/verl recorded 2026-06-04
~96,716 PyPI downloads last month
- https://github.com/volcengine/verl recorded 2026-06-04
adopter list incl. ByteDance, Qwen, Moonshot, NVIDIA research, Microsoft Research, Amazon, LinkedIn
Capability
5 high confidenceFrontier-tier within finetuning_code: broadest modern RL-algorithm coverage and demonstrated use training production frontier reasoning models. Calibrated alongside Megatron-LM (C5) as a scale/method definer for this category.
- https://github.com/volcengine/verl recorded 2026-06-04
PPO/GRPO/RLOO/GSPO/PRIME/DAPO/DrGRPO methods; HybridFlow; used for Seed-Thinking-v1.5 / Doubao
Unchanged since 2026-07-30 (last edited, not re-checked)