AI Potluck
Model components / Training & synthetic datasets

HH-RLHF (Anthropic Helpful/Harmless)

Anthropic

Anthropic's human preference dataset for RLHF, holding about 169,000 rows of chosen-versus-rejected response pairs on helpfulness and harmlessness alongside red-teaming dialogues. It was released with Anthropic's papers on training helpful and harmless assistants, and is intended for preference modeling rather than supervised dialogue training.

The Tulu 3 preference mixture does not use it. Verified 2026-08-13 via the Anthropic/hh-rlhf dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
mit
card
present
ungated
yes
rows
169,352
size
94.7MB

Publicly downloadable and ungated under the permissive MIT license.

Adoption

3 high confidence
3.0

32,927 downloads in the trailing 30 days for Anthropic/hh-rlhf, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Documented 2022 RLHF results and huge historical influence, but dropped from frontier open preference mixes (absent from Tulu 3 and OLMo 3 Dolci DPO).

Verified 2026-08-13