AI Potluck
Model components / Benchmark / eval datasets

Humanity's Last Exam (HLE)

Center for AI Safety

Multimodal frontier-knowledge benchmark from the Center for AI Safety and Scale AI, built with nearly a thousand subject-matter experts across more than 500 institutions. It holds 2,500 expert-written closed-ended questions, multiple-choice and short answer with text and images, across more than a hundred subjects, each with an unambiguous answer that is not quickly findable online.

A separate private held-out set detects overfitting, and it is described on the project site rather than the dataset card. Verified 2026-08-13 via the cais/hle dataset card on Hugging Face and lastexam.ai.

Openness

2 medium confidence
2.0
license
MIT(the public 2,500-question set)
questions+answers
public(2500-Q)
access
HF-gated(accept-conditions,no-redistribution,canary)
holdout
private

The public set is MIT-licensed and ships with its answers, but access is gated and a private held-out set exists, which is what decides the score. The private holdout is not mentioned on the Hugging Face card, so lastexam.ai is cited for it: it reads "We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting."

  • https://huggingface.co/datasets/cais/hle recorded 2026-08-13

    `license:mit` in the repo tags; the page says "You need to agree to share your contact information to access this dataset" and the embedded repo state reads `"gated":"auto"`; the card documents 2,500 questions in a single test split and the canary string "hle:3r2s:26b5c67b-..." that is "a superset of the BIG-bench canary string". The card makes no statement about a private held-out set

  • https://lastexam.ai recorded 2026-08-13

    "We publicly release these questions, while maintaining a private test set of held out questions to assess model overfitting", over a 2,500-question public set across more than a hundred subjects from nearly 1,000 expert contributors

Adoption

3 high confidence
3.0

34,715 Hugging Face downloads in the trailing 30 days for cais/hle, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models. HLE is a frontier-difficulty benchmark with flat, steady download volume.

Capability

5 high confidence
5.0

Far harder and broader than MMLU or MATH, with GPQA the nearest narrower peer on difficulty.

  • https://lastexam.ai recorded 2026-08-13

    nearly 1,000 subject-expert contributors across 500+ institutions; 2,500 questions across over a hundred subjects; "maintaining a private test set of held out questions to assess model overfitting"; the 04/03/2025 news entry finalizing the set at 2,500 questions

Verified 2026-08-12