EvalPlus
EvalPlus Team (UIUC)Extensions of HumanEval and MBPP that add roughly 80 times and 35 times more test cases through automated mutation, exposing models that pass weak tests while emitting buggy code. Published at NeurIPS 2023, with an EvalPerf companion for efficiency, and integrated with the bigcode evaluation harness.
The original benchmarks are saturated above 90 percent, which is why the Plus variants are used to separate frontier models. Verified 2026-08-13 via the evalplus/humanevalplus dataset card and the evalplus/evalplus repository.
Openness
5 high confidence- data
- downloadable(HF parquet, evalplus/humanevalplus + mbppplus)
- license
- Apache-2.0(OSI/open, on both framework and datasets)
- datasheet
- present(card documents structure: task_id/prompt/canonical_solution/test)
- redistributable
- yes
- framework-code
- open(Apache-2.0)
Clean open data: Apache-2.0 on both the datasets and the evaluation framework, downloadable, with a documented schema. Fully open.
- https://huggingface.co/datasets/evalplus/humanevalplus recorded 2026-08-13
`license:apache-2.0` in the repo tags; the embedded repo state reads `"gated":false`, so the parquet downloads without a barrier; the card documents the task_id/prompt/canonical_solution/test schema; 27,670 downloads in the trailing 30 days
- https://github.com/evalplus/evalplus recorded 2026-08-13
the repository's license tab still reads Apache-2.0, and the README still describes the HumanEval+ and MBPP+ variants and the bigcode-evaluation-harness integration
Adoption
3 medium confidence51,081 Hugging Face downloads in the trailing 30 days, summed across evalplus/humanevalplus and evalplus/mbppplus, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. These are the contamination-hardened HumanEval+ and MBPP+ variants, reported adopted by Meta and others.
- https://huggingface.co/api/datasets/evalplus/humanevalplus recorded 2026-08-12
27,086 downloads in the trailing 30 days for evalplus/humanevalplus
- https://huggingface.co/api/datasets/evalplus/mbppplus recorded 2026-08-12
23,995 downloads in the trailing 30 days for evalplus/mbppplus
Capability
not assessedA dataset is not capability-scored, so this axis is left unscored. Its distinguishing quality - far denser test suites that catch overfitting the base HumanEval and MBPP suites miss - is noted here rather than turned into a number.
- https://huggingface.co/datasets/evalplus/humanevalplus recorded 2026-08-13
a static problem corpus with denser tests; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-12