Humanity's Last Exam (HLE)
Center for AI SafetyA multi-modal frontier-knowledge benchmark from the Center for AI Safety and Scale AI, built with ~1,000 subject-matter experts: 2,500 expert-written closed-ended questions (MC + short answer, text+image) across 100+ subjects, each with an unambiguous answer not quickly findable online. Positioned as the hardest frontier academic benchmark.
The 2,500 public questions include answers under MIT, but the HF dataset is access-gated (accept conditions, no-redistribution + canary string) and a separate private held-out set detects overfitting -> scored gated (a license-weighted reading could argue open). ~35k monthly downloads.
Openness
2 medium confidence- questions+answers
- public(MIT,2500-Q)
- access
- HF-gated(accept-conditions,no-redistribution,canary)
- holdout
- private
Public set is MIT-licensed with answers, but access is gated and a private held-out set exists - judgment call (could be open under a license-weighted reading).
- https://huggingface.co/datasets/cais/hle recorded 2026-06-17
MIT, gated access, no-redistribution notice
Adoption
4 high confidence~35k monthly HF downloads; ~1.6k GitHub stars on the eval code.
- https://huggingface.co/datasets/cais/hle recorded 2026-06-17
~34,984 downloads/month
Capability
5 high confidenceFar harder and broader than MMLU/MATH; GPQA is a narrower difficulty peer.
- https://lastexam.ai recorded 2026-06-17
~1,000 experts, 2,500 Q, 100+ subjects, private holdout
Unchanged since 2026-06-17 (last edited, not re-checked)