AI Potluck
Model components / Benchmark / eval datasets

EvalPlus

EvalPlus Team (UIUC)

Extensions of HumanEval and MBPP that add roughly 80 times and 35 times more test cases through automated mutation, exposing models that pass weak tests while emitting buggy code. Published at NeurIPS 2023, with an EvalPerf companion for efficiency, and integrated with the bigcode evaluation harness.

The original benchmarks are saturated above 90 percent, which is why the Plus variants are used to separate frontier models. Verified 2026-08-13 via the evalplus/humanevalplus dataset card and the evalplus/evalplus repository.

Openness

5 high confidence
5.0
data
downloadable(HF parquet, evalplus/humanevalplus + mbppplus)
license
Apache-2.0(OSI/open, on both framework and datasets)
datasheet
present(card documents structure: task_id/prompt/canonical_solution/test)
redistributable
yes
framework-code
open(Apache-2.0)

Clean open data: Apache-2.0 on both the datasets and the evaluation framework, downloadable, with a documented schema. Fully open.

  • https://huggingface.co/datasets/evalplus/humanevalplus recorded 2026-08-13

    `license:apache-2.0` in the repo tags; the embedded repo state reads `"gated":false`, so the parquet downloads without a barrier; the card documents the task_id/prompt/canonical_solution/test schema; 27,670 downloads in the trailing 30 days

  • https://github.com/evalplus/evalplus recorded 2026-08-13

    the repository's license tab still reads Apache-2.0, and the README still describes the HumanEval+ and MBPP+ variants and the bigcode-evaluation-harness integration

Adoption

3 medium confidence
3.0

51,081 Hugging Face downloads in the trailing 30 days, summed across evalplus/humanevalplus and evalplus/mbppplus, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. These are the contamination-hardened HumanEval+ and MBPP+ variants, reported adopted by Meta and others.

Capability

not assessed

A dataset is not capability-scored, so this axis is left unscored. Its distinguishing quality - far denser test suites that catch overfitting the base HumanEval and MBPP suites miss - is noted here rather than turned into a number.

Verified 2026-08-12