MBPP
Google ResearchMostly Basic Python Problems, 974 crowd-sourced Python programming tasks designed to be solvable by entry-level programmers, each with a task description, solution, and three automated test cases. Introduced alongside Google's program synthesis work (2021). The complement to HumanEval: where HumanEval tests function-level generation with detailed docstrings, MBPP tests natural-language-to-code translation at a broader difficulty range. Widely used as a standard second benchmark alongside HumanEval.
MBPP (google-research-datasets/mbpp), ~974 crowd-sourced basic Python problems (full) + 427 hand-verified (sanitized); splits train/test/validation/prompt. From 'Program Synthesis with LLMs' (Austin et al., arXiv:2108.07732). Verified live on HF June 2026 (~198k downloads last month).
Openness
5 high confidence- license
- CC-BY-4.0(open data license)
- redistributability
- full(downloadable, ungated)
- datasheet
- yes(extensive dataset card)
- size
- ~974 full / 427 sanitized problems w/ test_list
- splits
- train/test/validation/prompt public
Fully open: CC-BY-4.0, ungated, downloadable, with an extensive dataset card. Clean open-data case; the standard companion to HumanEval for basic Python code-gen eval.
- https://huggingface.co/datasets/google-research-datasets/mbpp recorded 2026-06-04
CC-BY-4.0 license; not gated; ~974 full/427 sanitized problems; dataset card present; ~198k downloads last month
Adoption
4 medium confidenceStandard basic-Python code-generation benchmark, routinely paired with HumanEval and cited in code-model release reports. ~198k HF downloads/month on the canonical config; widely mirrored/forked. Annualized usage and citation-as-standard support a 4, just below the HumanEval/GSM8K tier.
- https://huggingface.co/datasets/google-research-datasets/mbpp recorded 2026-06-04
~198k downloads last month
Capability
not assessedDataset, not 'capable'. Largely saturated by frontier models; per recipe capability is n/a.
Unchanged since 2026-06-09 (last edited, not re-checked)