HH-RLHF (Anthropic Helpful/Harmless)
AnthropicAnthropic's human preference dataset for RLHF, holding about 169,000 rows of chosen-versus-rejected response pairs on helpfulness and harmlessness alongside red-teaming dialogues. It was released with Anthropic's papers on training helpful and harmless assistants, and is intended for preference modeling rather than supervised dialogue training.
The Tulu 3 preference mixture does not use it. Verified 2026-08-13 via the Anthropic/hh-rlhf dataset card on Hugging Face.
Openness
5 high confidence- license
- mit
- card
- present
- ungated
- yes
- rows
- 169,352
- size
- 94.7MB
Publicly downloadable and ungated under the permissive MIT license.
- https://huggingface.co/datasets/Anthropic/hh-rlhf recorded 2026-08-13
MIT license, ungated card, 169,352 rows, 94.7 MB
Adoption
3 high confidence32,927 downloads in the trailing 30 days for Anthropic/hh-rlhf, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/Anthropic/hh-rlhf recorded 2026-08-13
32,927 downloads in the trailing 30 days for Anthropic/hh-rlhf
Capability
3 high confidenceDocumented 2022 RLHF results and huge historical influence, but dropped from frontier open preference mixes (absent from Tulu 3 and OLMo 3 Dolci DPO).
- https://arxiv.org/abs/2204.05862 recorded 2026-08-13
Abstract - RLHF on the HH preference data improves performance on almost all NLP evaluations. The per-benchmark Elo, TruthfulQA and BBQ numbers are in the paper body, not on this page
- https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture recorded 2026-08-13
HH-RLHF is not used in the Tulu 3 preference mixture
Verified 2026-08-13