Humanity's Last Exam (HLE)
Center for AI SafetyMultimodal frontier-knowledge benchmark from the Center for AI Safety and Scale AI, built with nearly a thousand subject-matter experts across more than 500 institutions. It holds 2,500 expert-written closed-ended questions, multiple-choice and short answer with text and images, across more than a hundred subjects, each with an unambiguous answer that is not quickly findable online.
A separate private held-out set detects overfitting, and it is described on the project site rather than the dataset card. Verified 2026-08-13 via the cais/hle dataset card on Hugging Face and lastexam.ai.
Openness
2 medium confidence- license
- MIT(the public 2,500-question set)
- questions+answers
- public(2500-Q)
- access
- HF-gated(accept-conditions,no-redistribution,canary)
- holdout
- private
The public set is MIT-licensed and ships with its answers, but access is gated and a private held-out set exists, which is what decides the score. The private holdout is not mentioned on the Hugging Face card, so lastexam.ai is cited for it: it reads "We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting."
- https://huggingface.co/datasets/cais/hle recorded 2026-08-13
`license:mit` in the repo tags; the page says "You need to agree to share your contact information to access this dataset" and the embedded repo state reads `"gated":"auto"`; the card documents 2,500 questions in a single test split and the canary string "hle:3r2s:26b5c67b-..." that is "a superset of the BIG-bench canary string". The card makes no statement about a private held-out set
- https://lastexam.ai recorded 2026-08-13
"We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting", over a 2,500-question public set across more than a hundred subjects from nearly 1,000 expert contributors
Adoption
3 high confidence34,715 Hugging Face downloads in the trailing 30 days for cais/hle, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models. HLE is a frontier-difficulty benchmark with flat, steady download volume.
- https://huggingface.co/api/datasets/cais/hle recorded 2026-08-12
34,715 downloads in the trailing 30 days for cais/hle
Capability
5 high confidenceFar harder and broader than MMLU or MATH, with GPQA the nearest narrower peer on difficulty.
- https://lastexam.ai recorded 2026-08-13
nearly 1,000 subject-expert contributors across 500+ institutions; 2,500 questions across over a hundred subjects; "maintaining a private test set of held out questions to assess model overfitting"; the 04/03/2025 news entry finalizing the set at 2,500 questions
Verified 2026-08-12