AI Potluck
Model components / Fine-tuning code

OpenRLHF

OpenRLHF

Ray-based distributed RLHF/agentic-RL framework that distributes Actor, Reward, Reference, and Critic models across GPUs for full-scale 70B+ training. Picked over TRL when training scale exceeds a single node and the workload is RL-heavy (PPO, DAPO, REINFORCE++, GRPO, KTO), combines Ray + vLLM for high-throughput rollouts with DeepSpeed ZeRO-3 for memory-efficient training. 9.5K GitHub stars, 87 contributors, published research backing (arxiv 2405.11143); used by labs scaling preference optimization.

OpenRLHF (OpenRLHF/OpenRLHF), Apache-2.0, confirmed live June 2026 (~9.6K stars). Self-described 'first high-performance production-ready open source RLHF framework'; PPO, REINFORCE++, GRPO, RLOO; Ray scheduling + vLLM rollout + DeepSpeed; scales to 70B+ models; named users incl. Google, ByteDance, Tencent, Alibaba, Baidu, Allen AI.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(OpenRLHF/OpenRLHF)
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public, no proprietary tier.

Adoption

2 medium confidence
2.0

~9.6K GitHub stars; named adopters listed in repo (Google, ByteDance, Tencent, Alibaba, Baidu, China Telecom, Vivo, Allen AI, NexusFlow, Jülich). Specialist RLHF tool used by research/training teams rather than a mass-download library (no headline PyPI figure surfaced); level 2 on named-traction, not stars.

Capability

4 medium confidence
4.0

Strong distributed-RLHF capability (Ray+vLLM+DeepSpeed, 70B+), C4; below the Megatron-LM frontier-scale anchor and method scope is RL-centric rather than full-spectrum.

Unchanged since 2026-07-30 (last edited, not re-checked)