EvalPlus
EvalPlus Team (UIUC)HumanEval+ and MBPP+ extend the original benchmarks with 80x and 35x more test cases via automated mutation, exposing models that pass weak tests while emitting buggy code. Published at NeurIPS 2023; benchmarkers now prefer the Plus variants for discriminating among competitive frontier models, since the originals are saturated above 90%. OpenAI's o1 hits ~89% pass@1 on EvalPlus; ~1.7K GitHub stars and a maintained leaderboard make it the standard 'rigorous HumanEval' reference.
EvalPlus framework + datasets: HumanEval+ (164 problems, 80x more tests than original HumanEval) and MBPP+ (35x more tests than original MBPP), plus EvalPerf. Apache-2.0. HumanEval+ HF dataset verified live June 2026 (25,817 monthly downloads). Published NeurIPS 2023 (EvalPlus) / COLM 2024 (EvalPerf).
Openness
5 high confidence- data
- downloadable(HF parquet, evalplus/humanevalplus + mbppplus)
- license
- Apache-2.0(OSI/open, on both framework and datasets)
- datasheet
- present(card documents structure: task_id/prompt/canonical_solution/test)
- redistributable
- yes
- framework-code
- open(Apache-2.0)
Clean open data: Apache-2.0 on both datasets and the evaluation framework, downloadable with a documented schema, full open class.
- https://huggingface.co/datasets/evalplus/humanevalplus recorded 2026-06-04
Apache-2.0 license, 164 problems, 25,817 monthly downloads, documented schema
- https://github.com/evalplus/evalplus recorded 2026-06-04
Apache-2.0 framework; HumanEval+/MBPP+ described
Adoption
4 medium confidenceHumanEval+ alone: 25,817 monthly HF downloads (MBPP+ adds more). Top adoption signal for a benchmark = citation as a standard in frontier model release reports: EvalPlus is reported adopted by Meta (Llama 3.1/3.3), Alibaba (Qwen2.5-Coder/CodeQwen), DeepSeek, Allen AI (Tulu), Snowflake Arctic, StarCoder2. De facto standard for rigorous code-correctness eval. Level 4 on combined download volume + release-report standard status.
- https://huggingface.co/datasets/evalplus/humanevalplus recorded 2026-06-04
25,817 monthly downloads
- https://github.com/evalplus/evalplus recorded 2026-06-04
adopted by Meta/Qwen/DeepSeek/AllenAI/StarCoder2 for code eval; NeurIPS 2023
Capability
not assessedDataset, not capability-scored. Its distinguishing quality (far denser test suites that catch overfitting vs base HumanEval/MBPP) is a quality strength, noted but not on the 1-5 axis.
Unchanged since 2026-06-24 (last edited, not re-checked)