CodeContests
GoogleCompetitive-programming dataset built for the AlphaCode paper, holding 4,044 problems collected from Codeforces, CodeChef and other contest platforms, each with a description, correct and incorrect human solutions in several languages, and public, private and generated test cases.
An earlier version of this record stated 13,328 training problems; the dataset viewer shows 3,760 train and 4,044 total, and the record was corrected. Verified 2026-08-13 via the deepmind/code_contests dataset card on Hugging Face.
Openness
5 high confidence- license
- CC-BY-4.0(open data license, attribution-only)
- redistributability
- full(downloadable, ungated)
- datasheet
- yes(comprehensive dataset card + arXiv:2203.07814)
- size
- ~4k problems w/ rich test cases
- sources
- Codeforces/AtCoder/CodeChef/Aizu/HackerEarth + Description2Code(MIT)+CodeNet(Apache-2.0)
Fully open: CC-BY-4.0, ungated, downloadable, with a thorough datasheet documenting provenance. A clean open-data case.
- https://huggingface.co/datasets/deepmind/code_contests recorded 2026-08-13
`license:cc-by-4.0` in the repo tags; the embedded repo state reads `"gated":false`; the dataset card still renders; 51,142 downloads in the trailing 30 days
Adoption
3 medium confidence50,563 Hugging Face downloads in the trailing 30 days for deepmind/code_contests, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. This is the AlphaCode competitive-programming dataset, a recurring reference for code-reasoning evaluations but less ubiquitous than HumanEval or GSM8K in current release reports.
- https://huggingface.co/api/datasets/deepmind/code_contests recorded 2026-08-12
50,563 downloads in the trailing 30 days for deepmind/code_contests
Capability
not assessedA dataset is not 'capable', so this axis is left unscored. CodeContests is hard and discriminates well on code reasoning, but that is a quality strength rather than a capability score.
- https://huggingface.co/datasets/deepmind/code_contests recorded 2026-08-13
a static competitive-programming corpus of problems, solutions and test cases; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-12