IFEval
GoogleInstruction-Following Eval (IFEval) is a benchmark of 541 verifiable instructions used to measure how well models follow explicit constraints such as word-count limits, keyword inclusion, and format requirements. It is widely used for instruction-following and controllability evaluation.
IFEval (Instruction-Following Eval), 541 prompts with ~25 verifiable instruction types for programmatic (rule-based) scoring of LLM instruction-following. Paper arXiv:2311.07911. A core component of the HF Open LLM Leaderboard. HF dataset google/IFEval verified live June 2026.
Openness
5 high confidence- license
- Apache-2.0(OSI/open,redistributable)
- access
- public(not gated)
- datasheet
- present(card+structure+citation)
- splits
- public(single train split, no held-out)
Fully open: Apache-2.0, ungated, redistributable, with a complete dataset card.
- https://huggingface.co/datasets/google/IFEval recorded 2026-06-04
Apache-2.0 license, not gated, 541 examples, dataset card present
Adoption
4 high confidence~110,541 HF downloads in the trailing month (primary); a core benchmark on the HF Open LLM Leaderboard and routinely cited as the standard instruction-following eval, the top adoption signal for a benchmark.
- https://huggingface.co/datasets/google/IFEval recorded 2026-06-04
110,541 downloads last month; described as core Open LLM Leaderboard benchmark
Capability
not assessedDataset, not 'capable'; openness and adoption carry this category per recipe.
Unchanged since 2026-06-09 (last edited, not re-checked)