HH-RLHF (Anthropic Helpful/Harmless)
AnthropicHH-RLHF is Anthropic's human preference dataset for RLHF, containing ~169K rows of chosen-vs-rejected response pairs on helpfulness and harmlessness, plus red-teaming dialogues. Released alongside Anthropic's papers on training helpful and harmless assistants, it is intended for preference modeling rather than supervised dialogue training. It is one of the most influential early open RLHF preference datasets.
Verified live 2026-06-22 via primary sources. Publicly downloadable and ungated under the permissive MIT license.
Openness
5 high confidence- license
- mit
- card
- present
- ungated
- yes
- rows
- 169,352
- size
- 94.7MB
Publicly downloadable and ungated under the permissive MIT license.
- https://huggingface.co/datasets/Anthropic/hh-rlhf recorded 2026-06-22
MIT license, ungated card, 169,352 rows, 94.7 MB
Adoption
3 high confidence29,021 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/Anthropic/hh-rlhf recorded 2026-07-04
downloads field = 29,021 (30-day)
Capability
3 high confidenceDocumented 2022 RLHF results and huge historical influence, but dropped from frontier open preference mixes (absent from Tulu 3 and OLMo 3 Dolci DPO).
- https://arxiv.org/abs/2204.05862 recorded 2026-07-04
Anthropic RLHF on HH data improves helpfulness Elo, TruthfulQA, BBQ (self-reported)
- https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture recorded 2026-07-04
HH-RLHF is not used in the Tulu 3 preference mixture
Unchanged since 2026-07-04 (last edited, not re-checked)