AI Potluck
Model components / Training & synthetic datasets

UltraFeedback

OpenBMB

GPT-4-annotated preference dataset of roughly 64,000 prompts, 256,000 responses and 380,000 fine-grained feedback annotations across instruction-following, truthfulness, honesty and helpfulness. Prompts are drawn from six public datasets with responses from 17 different models. Built by OpenBMB, it is used to train reward and critic models in open RLHF and DPO pipelines.

Verified 2026-08-13 via the openbmb/UltraFeedback dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
mit
card
present
ungated
yes
rows
63,967
size
940MB

Publicly downloadable and ungated under the permissive MIT license.

Adoption

2 high confidence
2.0

5,456 downloads in the trailing 30 days for openbmb/UltraFeedback, which bands at 1K-10K, level 2 on the dataset adoption scale.

Capability

4 high confidence
4.0

Clean documented DPO before/after (Zephyr MT-Bench 6.64->7.34); its pipeline is the template for Tulu 3 / OLMo 2 preference data, with cleaned subsets still in use.

Verified 2026-08-13