AI Potluck
Model components / Benchmark / eval datasets

RewardBench

Allen Institute for AI

Benchmark for evaluating reward models, pairing prompts with chosen and rejected responses across chat, reasoning and safety subsets. RewardBench 2 extends it with harder, unseen human prompts, and both versions publish their datasets, results and a public leaderboard. The repository provides a common inference format for scoring most open reward models.

One product covering RewardBench and RewardBench 2; the repo README fronts both with datasets, results and leaderboard links. PyPI rewardbench's home_page points at a 404 path, so no pypi artifact is declared. Verified 2026-09-01 via the allenai/reward-bench GitHub API record, the LICENSE body, the repository README and both Hub dataset cards.

Openness

5 high confidence
5.0
license
ODC-BY(license odc-by on both Hub dataset cards
access
open(both Hub datasets ungated ("gated":false), downloadable parquet)
answers
released(chosen/rejected pairs are the labels and ship in the public splits
datasheet
present(dataset cards on both Hub releases with schema, subsets and citation)

ODC-BY on both dataset cards, ungated downloads, cards present. Ladder walk: license_tier open_data (odc-by is an enumerated open_data spelling) + documentation present → 5/open. Governing release is RewardBench 2, same license as v1, so the combine rule moves nothing.

Adoption

3 high confidence
3.0

14,234 Hugging Face downloads in the trailing 30 days summed across the two declared datasets (allenai/reward-bench 10,022 + allenai/reward-bench-2 4,212), the 10K-100K band, level 3 on the dataset scale. RewardBench is the standard reward-model leaderboard, but the band rests on the measured figure.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored, mirroring the category pattern (gsm8k).

Verified 2026-09-01