UltraFeedback
OpenBMBGPT-4-annotated preference dataset of roughly 64,000 prompts, 256,000 responses and 380,000 fine-grained feedback annotations across instruction-following, truthfulness, honesty and helpfulness. Prompts are drawn from six public datasets with responses from 17 different models. Built by OpenBMB, it is used to train reward and critic models in open RLHF and DPO pipelines.
Verified 2026-08-13 via the openbmb/UltraFeedback dataset card on Hugging Face.
Openness
5 high confidence- license
- mit
- card
- present
- ungated
- yes
- rows
- 63,967
- size
- 940MB
Publicly downloadable and ungated under the permissive MIT license.
- https://huggingface.co/datasets/openbmb/UltraFeedback recorded 2026-08-13
MIT license, ungated card, 63,967 rows / 940 MB
Adoption
2 high confidence5,456 downloads in the trailing 30 days for openbmb/UltraFeedback, which bands at 1K-10K, level 2 on the dataset adoption scale.
- https://huggingface.co/api/datasets/openbmb/UltraFeedback recorded 2026-08-13
5,456 downloads in the trailing 30 days for openbmb/UltraFeedback
Capability
4 high confidenceClean documented DPO before/after (Zephyr MT-Bench 6.64->7.34); its pipeline is the template for Tulu 3 / OLMo 2 preference data, with cleaned subsets still in use.
- https://ar5iv.labs.arxiv.org/html/2310.16944 recorded 2026-08-13
Zephyr: DPO on ultrafeedback_binarized lifts MT-Bench 6.64->7.00, final model 7.34
- https://arxiv.org/abs/2411.15124 recorded 2026-08-13
Tulu 3 releases an open preference-data pipeline at scale; the UltraFeedback lineage is described in the paper body, not on this page
Verified 2026-08-13