MBPP
Google ResearchMostly Basic Python Problems: 974 crowd-sourced Python tasks solvable by an entry-level programmer, each with a task description, a reference solution and three automated test cases, plus a 427-problem hand-verified subset. Where HumanEval tests function-level generation from a detailed docstring, MBPP tests natural-language-to-code translation across a broader difficulty range.
Verified 2026-08-13 via the google-research-datasets/mbpp dataset card on Hugging Face.
Openness
5 high confidence- license
- CC-BY-4.0(open data license)
- redistributability
- full(downloadable, ungated)
- datasheet
- yes(extensive dataset card)
- size
- ~974 full / 427 sanitized problems w/ test_list
- splits
- train/test/validation/prompt public
Fully open: CC-BY-4.0, ungated, downloadable, with an extensive dataset card. Clean open-data case; the standard companion to HumanEval for basic Python code-gen eval.
- https://huggingface.co/datasets/google-research-datasets/mbpp recorded 2026-08-13
`license:cc-by-4.0` in the repo tags; the embedded repo state reads `"gated":false`; the dataset card still renders with the full and sanitized configs; 217,014 downloads in the trailing 30 days
Adoption
4 medium confidence208,317 Hugging Face downloads in the trailing 30 days for google-research-datasets/mbpp, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. MBPP is a standard basic-Python code-generation benchmark, routinely paired with HumanEval in code-model release reports.
- https://huggingface.co/api/datasets/google-research-datasets/mbpp recorded 2026-08-12
208,317 downloads in the trailing 30 days for google-research-datasets/mbpp
Capability
not assessedA dataset is not 'capable', so this axis is left unscored. MBPP is largely saturated by frontier models, but that is a quality caveat rather than a capability score.
- https://huggingface.co/datasets/google-research-datasets/mbpp recorded 2026-08-13
a static corpus of basic Python tasks with text, code and test_list columns; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-12