AI Potluck
Model components / Training & synthetic datasets

UltraFeedback

OpenBMB

UltraFeedback is a large-scale GPT-4-annotated preference dataset comprising ~64K prompts, ~256K responses, and ~380K fine-grained feedback annotations across instruction-following, truthfulness, honesty, and helpfulness. Prompts are drawn from six public datasets with responses from 17 different models. Built by OpenBMB, it is a foundational dataset for training reward and critic models and underpins many open RLHF/DPO pipelines.

Verified live 2026-06-22 via primary sources. Publicly downloadable and ungated under the permissive MIT license.

Openness

5 high confidence
5.0
license
mit
card
present
ungated
yes
rows
63,967
size
940MB

Publicly downloadable and ungated under the permissive MIT license.

Adoption

2 high confidence
2.0

4,496 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 high confidence
4.0

Clean documented DPO before/after (Zephyr MT-Bench 6.64->7.34); its pipeline is the template for Tulu 3 / OLMo 2 preference data, with cleaned subsets still in use.

Unchanged since 2026-07-04 (last edited, not re-checked)