AI Potluck
Model components / Benchmark / eval datasets

MBPP

Google Research

Mostly Basic Python Problems, 974 crowd-sourced Python programming tasks designed to be solvable by entry-level programmers, each with a task description, solution, and three automated test cases. Introduced alongside Google's program synthesis work (2021). The complement to HumanEval: where HumanEval tests function-level generation with detailed docstrings, MBPP tests natural-language-to-code translation at a broader difficulty range. Widely used as a standard second benchmark alongside HumanEval.

MBPP (google-research-datasets/mbpp), ~974 crowd-sourced basic Python problems (full) + 427 hand-verified (sanitized); splits train/test/validation/prompt. From 'Program Synthesis with LLMs' (Austin et al., arXiv:2108.07732). Verified live on HF June 2026 (~198k downloads last month).

Openness

5 high confidence
5.0
license
CC-BY-4.0(open data license)
redistributability
full(downloadable, ungated)
datasheet
yes(extensive dataset card)
size
~974 full / 427 sanitized problems w/ test_list
splits
train/test/validation/prompt public

Fully open: CC-BY-4.0, ungated, downloadable, with an extensive dataset card. Clean open-data case; the standard companion to HumanEval for basic Python code-gen eval.

Adoption

4 medium confidence
4.0

Standard basic-Python code-generation benchmark, routinely paired with HumanEval and cited in code-model release reports. ~198k HF downloads/month on the canonical config; widely mirrored/forked. Annualized usage and citation-as-standard support a 4, just below the HumanEval/GSM8K tier.

Capability

not assessed

Dataset, not 'capable'. Largely saturated by frontier models; per recipe capability is n/a.

Unchanged since 2026-06-09 (last edited, not re-checked)