Epoch AI Benchmarks
Epoch AIIndependent benchmarking lab that runs and maintains canonical leaderboards for high-stakes evals (SWE-bench Verified, FrontierMath, GPQA Diamond) and operates the FrontierMath private math benchmark. Picked when researchers want the authoritative scoreboard for a specific benchmark; Epoch's SWE-bench Verified leaderboard is the one frontier labs reference in announcements. Funded by Open Philanthropy and lab partners.
Epoch AI benchmark database + public dashboard scoring frontier models on hard tasks (FrontierMath, GPQA Diamond, SWE-bench Verified, SimpleQA Verified, MATH Level 5, OTIS Mock AIME, Chess Puzzles). Verified live June 2026; org github.com/epoch-research has MIT/Apache eval repos (SWE-bench docker registry, eci-public, MirrorCode-data), but FrontierMath dataset is held private.
Openness
2 medium confidence- dashboard
- public(free to browse)
- eval-code
- partial-open(several org repos MIT/Apache - SWE-bench docker registry, eci-public, MirrorCode-data)
- benchmark-data
- mixed(FrontierMath gated/private to prevent contamination
Mixed-tier: public results dashboard plus some MIT/Apache-licensed eval components, but the headline benchmark (FrontierMath) is deliberately private and there is no single OSI-licensed end-to-end harness, closer to source_available than fully open_source.
- https://github.com/epoch-research recorded 2026-06-04
public org repos incl SWE-bench docker registry, eci-public (MIT), MirrorCode-data; mostly MIT/Apache-2.0
- https://epoch.ai/benchmarks recorded 2026-06-04
public benchmark dashboard running FrontierMath/GPQA Diamond/SWE-bench Verified/SimpleQA Verified across frontier models
Adoption
3 medium confidenceEpoch's benchmark results (esp. FrontierMath) are a widely-referenced capability signal cited across frontier-model coverage and research; the public dashboard is a recognized trends source. No standalone usage/traffic count published, so reported_traction; level 3 reflects broad authoritative reference rather than a measured figure. No stars-only fallback used.
- https://epoch.ai/benchmarks recorded 2026-06-04
public database of leading-model benchmark results positioned as authoritative AI-capability tracker
Capability
4 medium confidenceStrong feature_matrix on hard/frontier benchmark coverage and trusted standardized results, especially math. Not a general agentic-eval harness; held-out design limits self-serve reproducibility. Placed at 4, authoritative but not a frontier-defining general harness like lm-eval-harness.
- https://epoch.ai/benchmarks recorded 2026-06-04
operated benchmarks list: FrontierMath, GPQA Diamond, SWE-bench Verified, SimpleQA Verified, MATH Level 5, OTIS Mock AIME, Chess Puzzles
Unchanged since 2026-06-24 (last edited, not re-checked)