SWE-bench Verified
Princeton NLP / OpenAI500 human-validated instances drawn from the wider SWE-bench dataset, each pairing a real GitHub issue with a test-verified solution patch and the fail-to-pass and pass-to-pass test lists that grade it. It was curated by OpenAI from the SWE-bench test set to reduce noise in the original 2,294-instance set.
The dataset card declares no license, and the harness repository's README points at a LICENSE.md path that does not exist. Verified 2026-08-13 via the SWE-bench_Verified card, its raw README and the SWE-bench repository LICENSE.
Openness
5 medium confidence- license
- MIT(SWE-bench repository states MIT over its code and data
- redistributability
- full(downloadable parquet, ungated)
- datasheet
- yes(dataset card + arXiv:2310.06770)
- held-out
- no hidden split in Verified itself(the public 500-instance set is fully released
- data-source
- real GitHub PR/issue pairs
Ungated and downloadable, with full gold patches and tests, a datasheet present, and the license stated. The canonical SWE-bench repository describes itself as "Code and data for the following works" and lists SWE-bench among them; its "Citation & license" section says "MIT license", with the citation block directly beneath headed "For SWE-bench (Verified)". The Hugging Face card for Verified declares no license field, and where a repository license and a distribution-point license differ the repository license governs, so MIT is what applies and this counts as fully open. (Distinct from the broader and hidden SWE-bench variants, which can carry held-out splits.)
- https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified recorded 2026-08-13
the embedded repo state reads `"gated":false` and the tags array carries no license tag; the dataset card renders over the 500-instance test split with the gold patch and the FAIL_TO_PASS/PASS_TO_PASS test lists; 413,323 downloads in the trailing 30 days
- https://github.com/SWE-bench/SWE-bench/blob/main/LICENSE recorded 2026-06-04
MIT License for the SWE-bench harness
- https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified/blob/main/README.md recorded 2026-06-04
dataset card metadata has no license field declared
- https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified/raw/main/README.md recorded 2026-08-13
card front matter carries dataset_info, splits and configs and no `license:` key; body describes the 500 human-validated instances and the swebench.com leaderboard, so the data is unlicensed at source. The repository holds only .gitattributes, README.md and one parquet file, so there is no LICENSE file alongside the data either
- https://raw.githubusercontent.com/SWE-bench/SWE-bench/main/README.md recorded 2026-08-13
"Code and data for the following works:" listing SWE-bench and SWE-bench Multimodal; the "Citation & license" section reads "MIT license. Check `LICENSE.md`." and the citation immediately under it is headed "For SWE-bench (Verified)"; the news entry for Aug. 13, 2024 introduces Verified as a 500-problem subset produced with OpenAI Preparedness
- https://raw.githubusercontent.com/SWE-bench/SWE-bench/main/LICENSE recorded 2026-08-13
MIT License, copyright 2023 the SWE-bench authors. Note the README points at `LICENSE.md`, which 404s; the file is `LICENSE`
Adoption
4 high confidenceThe de-facto standard agentic-coding benchmark: it is cited as the headline coding metric in nearly every frontier model release - GPT-5, Claude Opus, DeepSeek, Qwen and Kimi all report SWE-bench Verified - alongside a live, heavily-submitted leaderboard at swebench.com. Hugging Face records 413,323 downloads in the trailing 30 days across the declared artifacts, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. It is the strongest citation-as-standard signal in this category.
- https://huggingface.co/api/datasets/princeton-nlp/SWE-bench_Verified recorded 2026-08-13
413,323 downloads in the trailing 30 days for princeton-nlp/SWE-bench_Verified
Capability
not assessedA dataset is not 'capable', so this axis is left unscored. SWE-bench Verified is notably contamination-resistant and discriminating for coding agents, since it runs against real repository state and executable tests, but that quality is noted here rather than turned into a number.
- https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified recorded 2026-08-13
a static 500-instance corpus of repo state, issue text, gold patch and tests; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-13