AI Potluck
Model components / Benchmark / eval datasets

EvalPlus

EvalPlus Team (UIUC)

HumanEval+ and MBPP+ extend the original benchmarks with 80x and 35x more test cases via automated mutation, exposing models that pass weak tests while emitting buggy code. Published at NeurIPS 2023; benchmarkers now prefer the Plus variants for discriminating among competitive frontier models, since the originals are saturated above 90%. OpenAI's o1 hits ~89% pass@1 on EvalPlus; ~1.7K GitHub stars and a maintained leaderboard make it the standard 'rigorous HumanEval' reference.

EvalPlus framework + datasets: HumanEval+ (164 problems, 80x more tests than original HumanEval) and MBPP+ (35x more tests than original MBPP), plus EvalPerf. Apache-2.0. HumanEval+ HF dataset verified live June 2026 (25,817 monthly downloads). Published NeurIPS 2023 (EvalPlus) / COLM 2024 (EvalPerf).

Openness

5 high confidence
5.0
data
downloadable(HF parquet, evalplus/humanevalplus + mbppplus)
license
Apache-2.0(OSI/open, on both framework and datasets)
datasheet
present(card documents structure: task_id/prompt/canonical_solution/test)
redistributable
yes
framework-code
open(Apache-2.0)

Clean open data: Apache-2.0 on both datasets and the evaluation framework, downloadable with a documented schema, full open class.

Adoption

4 medium confidence
4.0

HumanEval+ alone: 25,817 monthly HF downloads (MBPP+ adds more). Top adoption signal for a benchmark = citation as a standard in frontier model release reports: EvalPlus is reported adopted by Meta (Llama 3.1/3.3), Alibaba (Qwen2.5-Coder/CodeQwen), DeepSeek, Allen AI (Tulu), Snowflake Arctic, StarCoder2. De facto standard for rigorous code-correctness eval. Level 4 on combined download volume + release-report standard status.

Capability

not assessed

Dataset, not capability-scored. Its distinguishing quality (far denser test suites that catch overfitting vs base HumanEval/MBPP) is a quality strength, noted but not on the 1-5 axis.

Unchanged since 2026-06-24 (last edited, not re-checked)