Epoch AI Benchmarks
Epoch AIIndependent benchmarking lab that runs and maintains leaderboards for high-stakes evaluations, including SWE-bench Verified, GPQA Diamond, SimpleQA Verified and MATH Level 5, and operates its own FrontierMath benchmark across four difficulty tiers. Its public dashboard scores frontier models on all of them.
The FrontierMath problems are held back by design, so the benchmark behind the dashboard cannot be re-run externally. Verified 2026-08-13 via the Epoch benchmarks dashboard, the FrontierMath page and the organization's repository listing.
Openness
2 medium confidence- dashboard
- public(free to browse)
- eval-code
- partial-open(several org repos MIT/Apache - SWE-bench docker registry, eci-public, MirrorCode-data)
- benchmark-data
- mixed(FrontierMath gated/private to prevent contamination
- source
- partial(public dashboard and MIT component repos, but no end-to-end harness published and the FrontierMath problems are unpublished)
Mixed: a public results dashboard plus some MIT- and Apache-licensed eval components, but the headline benchmark, FrontierMath, is deliberately private, and there is no single OSI-licensed end-to-end harness. That puts it closer to source-available than to fully open source.
- https://github.com/epoch-research recorded 2026-06-04
public org repos incl SWE-bench docker registry, eci-public (MIT), MirrorCode-data; mostly MIT/Apache-2.0
- https://epoch.ai/benchmarks recorded 2026-06-04
public benchmark dashboard running FrontierMath/GPQA Diamond/SWE-bench Verified/SimpleQA Verified across frontier models
- https://epoch.ai/frontiermath recorded 2026-08-11
Epoch describes FrontierMath Tiers 1-4 as 'a benchmark of several hundred unpublished, highly challenging mathematics problems'. The problems are held back by design, so the benchmark behind the public dashboard cannot be run by a reader.
- https://api.github.com/orgs/epoch-research/repos recorded 2026-08-13
33 public repos in the epoch-research org. Licensed eval components include SWE-bench and SWE-bench-c7d4d9d4 (MIT), eci-public (MIT), MirrorCode (MIT), benchmark-stitching (MIT), epochutils (MIT), ftm (MIT) and timelines-direct-approach (Apache-2.0); MirrorCode-data and many others carry no license at all. There is no single end-to-end harness repo among them.
- https://epoch.ai/frontiermath recorded 2026-08-13
Still 'A benchmark of several hundred unpublished, highly challenging mathematics problems', with Tiers 1-3 and Tier 4 deliberately withheld, so the headline benchmark behind the dashboard cannot be run by a reader.
Adoption
3 medium confidenceEpoch's benchmark results, FrontierMath above all, are a widely referenced capability signal cited across frontier-model coverage and research, and the public dashboard is a recognized source for trends. It is free to browse and spans roughly 60 benchmarks across Epoch-run, creator-run and developer-run sets, none of it accompanied by a visitor or usage count. With no standalone usage or traffic figure published anywhere, level 3 reflects broad authoritative reference rather than a measured number, and a GitHub star count was not used as a substitute.
- https://epoch.ai/benchmarks recorded 2026-06-04
public database of leading-model benchmark results positioned as authoritative AI-capability tracker
- https://epoch.ai/benchmarks recorded 2026-08-13
Free public dashboard letting a reader 'explore trends in AI capabilities across time, by benchmark, or by model', spanning roughly 60 benchmarks in three groups. No traffic, visitor or usage figure is published on it.
Capability
4 medium confidenceStrong on hard, frontier benchmark coverage and trusted standardized results, especially in math. It is not a general agentic-eval harness, and the held-out design limits self-serve reproducibility, so it lands at 4 - authoritative, but not a frontier-defining general harness in the way lm-evaluation-harness is. The Epoch-run set is wider than the list recorded here: eleven benchmarks, adding MirrorCode, Earthborne Rangers and Mystery Game Puzzles, and splitting FrontierMath into Tier 4 (v2) and Tiers 1-3 (v2).
- https://epoch.ai/benchmarks recorded 2026-06-04
operated benchmarks list: FrontierMath, GPQA Diamond, SWE-bench Verified, SimpleQA Verified, MATH Level 5, OTIS Mock AIME, Chess Puzzles
- https://epoch.ai/benchmarks recorded 2026-08-13
Epoch AI-run benchmarks listed on the page: FrontierMath Tier 4 (v2), FrontierMath Tiers 1-3 (v2), SWE-bench Verified, MirrorCode, Chess Puzzles, Earthborne Rangers (EBR-bench), Mystery Game Puzzles, SimpleQA Verified, GPQA Diamond, OTIS Mock AIME 2024-2025 and MATH Level 5, plus an expansion of FrontierMath into 50 unsolved research problems.
Verified 2026-08-13