AI Potluck
Model components / Benchmark / eval datasets

IFEval

Google

Instruction-Following Eval (IFEval) is a benchmark of 541 verifiable instructions used to measure how well models follow explicit constraints such as word-count limits, keyword inclusion, and format requirements. It is widely used for instruction-following and controllability evaluation.

IFEval (Instruction-Following Eval), 541 prompts with ~25 verifiable instruction types for programmatic (rule-based) scoring of LLM instruction-following. Paper arXiv:2311.07911. A core component of the HF Open LLM Leaderboard. HF dataset google/IFEval verified live June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI/open,redistributable)
access
public(not gated)
datasheet
present(card+structure+citation)
splits
public(single train split, no held-out)

Fully open: Apache-2.0, ungated, redistributable, with a complete dataset card.

Adoption

4 high confidence
4.0

~110,541 HF downloads in the trailing month (primary); a core benchmark on the HF Open LLM Leaderboard and routinely cited as the standard instruction-following eval, the top adoption signal for a benchmark.

Capability

not assessed

Dataset, not 'capable'; openness and adoption carry this category per recipe.

Unchanged since 2026-06-09 (last edited, not re-checked)