TRL
Hugging FaceThe de facto reference library for post-training transformers, supports SFT, DPO, GRPO, PPO, KTO, reward modeling, and other preference-optimization methods through unified Trainer APIs. Picked over alternatives because it sits at the center of the HF ecosystem (transformers, datasets, accelerate, PEFT) and ships canonical reference implementations that papers and other frameworks build on. 18.4K GitHub stars, 482 contributors, ~400 commits in 90 days; extremely active.
TRL (Transformer Reinforcement Learning) post-training library; repo huggingface/trl active, LICENSE confirmed live June 2026. Supplies SFT, DPO/IPO/KTO/ORPO, PPO, GRPO trainers built on HF Transformers/Accelerate; the de-facto reference post-training/RLHF library in the HF ecosystem.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(huggingface/trl)
- governance
- Hugging Face
- managed-tier
- none
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public, no proprietary core or managed training SaaS gate.
- https://github.com/huggingface/trl/blob/main/LICENSE recorded 2026-06-04
Apache License Version 2.0 text
Adoption
4 high confidence~3.81M PyPI downloads last month; standard library for RLHF/post-training in the HF stack, integrated with Transformers/Accelerate and used across labs and open-model post-training pipelines.
- https://pypistats.org/packages/trl recorded 2026-06-04
Downloads last month: 3,809,112
Capability
4 high confidenceBroadest post-training method coverage among the open libraries; one tier below the Megatron-LM frontier-scale-training anchor on demonstrated large-scale parallelism.
- https://github.com/huggingface/trl/blob/main/LICENSE recorded 2026-06-04
huggingface/trl repo (TRL post-training library) confirmed live
Unchanged since 2026-07-30 (last edited, not re-checked)