AI Potluck
Model components / Training & synthetic datasets

HH-RLHF (Anthropic Helpful/Harmless)

Anthropic

HH-RLHF is Anthropic's human preference dataset for RLHF, containing ~169K rows of chosen-vs-rejected response pairs on helpfulness and harmlessness, plus red-teaming dialogues. Released alongside Anthropic's papers on training helpful and harmless assistants, it is intended for preference modeling rather than supervised dialogue training. It is one of the most influential early open RLHF preference datasets.

Verified live 2026-06-22 via primary sources. Publicly downloadable and ungated under the permissive MIT license.

Openness

5 high confidence
5.0
license
mit
card
present
ungated
yes
rows
169,352
size
94.7MB

Publicly downloadable and ungated under the permissive MIT license.

Adoption

3 high confidence
3.0

29,021 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Documented 2022 RLHF results and huge historical influence, but dropped from frontier open preference mixes (absent from Tulu 3 and OLMo 3 Dolci DPO).

Unchanged since 2026-07-04 (last edited, not re-checked)