UltraFeedback
OpenBMBUltraFeedback is a large-scale GPT-4-annotated preference dataset comprising ~64K prompts, ~256K responses, and ~380K fine-grained feedback annotations across instruction-following, truthfulness, honesty, and helpfulness. Prompts are drawn from six public datasets with responses from 17 different models. Built by OpenBMB, it is a foundational dataset for training reward and critic models and underpins many open RLHF/DPO pipelines.
Verified live 2026-06-22 via primary sources. Publicly downloadable and ungated under the permissive MIT license.
Openness
5 high confidence- license
- mit
- card
- present
- ungated
- yes
- rows
- 63,967
- size
- 940MB
Publicly downloadable and ungated under the permissive MIT license.
- https://huggingface.co/datasets/openbmb/UltraFeedback recorded 2026-06-22
MIT license, ungated card, 63,967 rows / 940 MB
Adoption
2 high confidence4,496 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/openbmb/UltraFeedback recorded 2026-07-04
downloads field = 4,496 (30-day)
Capability
4 high confidenceClean documented DPO before/after (Zephyr MT-Bench 6.64->7.34); its pipeline is the template for Tulu 3 / OLMo 2 preference data, with cleaned subsets still in use.
- https://ar5iv.labs.arxiv.org/html/2310.16944 recorded 2026-07-04
Zephyr: DPO on ultrafeedback_binarized lifts MT-Bench 6.64->7.00, final model 7.34
- https://arxiv.org/abs/2411.15124 recorded 2026-07-04
Tulu 3 extends the UltraFeedback pipeline to scale preference data
Unchanged since 2026-07-04 (last edited, not re-checked)