AI Potluck
Model components / Benchmark / eval datasets

Humanity's Last Exam (HLE)

Center for AI Safety

A multi-modal frontier-knowledge benchmark from the Center for AI Safety and Scale AI, built with ~1,000 subject-matter experts: 2,500 expert-written closed-ended questions (MC + short answer, text+image) across 100+ subjects, each with an unambiguous answer not quickly findable online. Positioned as the hardest frontier academic benchmark.

The 2,500 public questions include answers under MIT, but the HF dataset is access-gated (accept conditions, no-redistribution + canary string) and a separate private held-out set detects overfitting -> scored gated (a license-weighted reading could argue open). ~35k monthly downloads.

Openness

2 medium confidence
2.0
questions+answers
public(MIT,2500-Q)
access
HF-gated(accept-conditions,no-redistribution,canary)
holdout
private

Public set is MIT-licensed with answers, but access is gated and a private held-out set exists - judgment call (could be open under a license-weighted reading).

Adoption

4 high confidence
4.0

~35k monthly HF downloads; ~1.6k GitHub stars on the eval code.

Capability

5 high confidence
5.0

Far harder and broader than MMLU/MATH; GPQA is a narrower difficulty peer.

  • https://lastexam.ai recorded 2026-06-17

    ~1,000 experts, 2,500 Q, 100+ subjects, private holdout

Unchanged since 2026-06-17 (last edited, not re-checked)