AI Potluck
Model components / Fine-tuning code

NeMo-RL

NVIDIA

NVIDIA's scalable open source toolkit for RL post-training, successor to NeMo-Aligner. Supports DPO, PPO, GRPO, and reward modeling on Megatron-LM and FSDP backends with first-class multi-node H100/B200 scaling. Picked when teams are already on the NVIDIA NeMo stack and need RL post-training that scales to thousands of GPUs. 1.6K GitHub stars but 181 commits in 90 days and a growing contributor base, well-resourced; pairs with NeMo Customizer (managed) for enterprise users.

NVIDIA NeMo-RL, v0.6.0 released Apr 30 2026. Scalable post-training library for LLMs/VLMs (small-scale to multi-node). Confirmed live on GitHub June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
maintainer
NVIDIA
core-gated
ungated

Apache-2.0, full source public; GPU orientation is a hardware practicality, not a license restriction.

Adoption

2 medium confidence
2.0

No PyPI download or named-user volume located this run; only ~1.7k GitHub stars (last-resort signal, caps at level 3). Young (v0.6.0, Apr 2026) and narrower than the broader NeMo framework. Level 2 on stars_fallback; honest signal is limited.

Capability

5 high confidence
5.0

Frontier-tier scale/parallelism within finetuning_code: Megatron-Core backend with 6D parallelism and modern RL/preference method coverage. Capability (does it train at scale, well?) is high even though adoption is still small; calibrated as C5 alongside Megatron-LM and verl.

Unchanged since 2026-07-30 (last edited, not re-checked)