HumanEval
OpenAI164 hand-written Python programming problems with function signatures, docstrings, and unit tests. The original code generation benchmark, introduced in the Codex paper (2021). Still the most widely cited code eval dataset; virtually every code model paper reports pass@k on HumanEval, making it the de facto baseline despite its small size and Python-only scope.
HumanEval (openai/openai_humaneval), 164 hand-written Python problems (prompt + canonical solution + unit tests), single test split, ~92 kB. From 'Evaluating LLMs Trained on Code' (Codex paper, arXiv:2107.03374). Verified live on HF June 2026 (~287k downloads last month).
Openness
5 high confidence- license
- MIT(OSI-style permissive, fully redistributable)
- redistributability
- full(downloadable parquet, no gating)
- datasheet
- yes(dataset card)
- size
- 164 problems/test split
- contamination
- hand-written to avoid GitHub-dump overlap (per design)
Fully open: MIT license, ungated, downloadable, with a dataset card. Textbook open-data case for this category.
- https://huggingface.co/datasets/openai/openai_humaneval recorded 2026-06-04
MIT license; not gated; 164 problems; dataset card present; ~287k downloads last month
Adoption
4 high confidenceFoundational code-generation benchmark cited as a standard in code-model release reports for years (pass@k metric originated here). ~287k HF downloads/month on the canonical config alone, plus heavy mirrored/forked use; sustained 1M+/yr equivalent. The reference code-gen eval despite being partially saturated by frontier models.
- https://huggingface.co/datasets/openai/openai_humaneval recorded 2026-06-04
~287k downloads last month on canonical HumanEval config
Capability
not assessedA dataset is not 'capable'. Worth noting it is now largely saturated/contaminated as a discriminator, but per recipe capability is n/a here.
Unchanged since 2026-06-09 (last edited, not re-checked)