MLE-bench
OpenAIBenchmark measuring how well AI agents perform machine learning engineering, built from 75 Kaggle competitions. OpenAI ships the preparation and grading code plus new train/test splits, and the raw competition data is downloaded from Kaggle under each competition's own rules. A public leaderboard tracks agent submissions, though new submissions are paused as of April 2026.
The LICENSE is MIT over the code with an explicit carve-out excluding the external datasets, and the corpus is downloaded from Kaggle with credentials under each competition's own rules. GitHub-only channels; no first-party HF dataset or PyPI package. Verified 2026-09-01 via the openai/mle-bench GitHub API record, the LICENSE body and the repository README.
Openness
3 high confidence- license
- per-component(one component per Kaggle competition: the repo LICENSE is MIT over the code only, with the explicit NOTE "This license applies to the code in this repository, but not the external datasets and files that may be downloaded while using this package", so the corpus defers to each competition's own rules)
- access
- terms-acceptance(data downloaded via the Kaggle API with Kaggle credentials, under Kaggle's per-competition rules
- answers
- released(the benchmark's own test splits are constructed by the prepare scripts from public training sets, and grading scripts ship in the repo — no split hidden by MLE-bench itself)
- datasheet
- yes(repository README documents composition (75 competitions), preparation and grading
- arXiv
- 2410.07095)
The gate decides this one. The corpus is obtainable only through Kaggle with an account and per-competition terms, so availability is terms-acceptance and the gate rung fires at 3/gated before the deferred-license rungs are reached; no license tier can lift a gated corpus past 3. The MIT covers only the harness code, and the LICENSE body says so explicitly — which is also why GitHub reports NOASSERTION.
- https://raw.githubusercontent.com/openai/mle-bench/main/LICENSE recorded 2026-09-01
MIT License, Copyright (c) 2024 OpenAI, standard text, then: "NOTE: This license applies to the code in this repository, but not the external datasets and files that may be downloaded while using this package." — the carve-out that makes the corpus deferred and explains GitHub's NOASSERTION
- https://raw.githubusercontent.com/openai/mle-bench/main/README.md recorded 2026-09-01
dataset is 75 Kaggle competitions; prepare scripts split public training sets into new train/test; raw data downloaded with the Kaggle API requiring kaggle.json credentials; grading scripts per competition; leaderboard paused for new submissions 2026-04-24
- https://api.github.com/repos/openai/mle-bench recorded 2026-09-01
license spdx NOASSERTION (the classifier choking on the carve-out note), 1,727 stars, homepage openai.com/index/mle-bench
Adoption
2 medium confidenceGitHub is the only first-party channel — no HF dataset, no PyPI package, and the data itself is fetched from Kaggle where no per-benchmark download figure is published. The honest call is the stars fallback: 1,727 stars on openai/mle-bench, the 1K-10K band, level 2, capped at 3 by the rubric. Not banded as reported_traction because no vendor traction claim with a figure exists to cite.
- https://api.github.com/repos/openai/mle-bench recorded 2026-09-01
stargazers_count 1727
Capability
not assessedA dataset is not 'capable', so this axis is left unscored, mirroring the category pattern (gsm8k).
- https://raw.githubusercontent.com/openai/mle-bench/main/README.md recorded 2026-09-01
a benchmark definition plus leaderboard of OTHER systems' scores; no capability claim about the artifact itself
Verified 2026-09-01