AI Potluck
Organization

Epoch AI

lab

Openness profile

1 product on the map — 1 open-ish.

Epoch AI Benchmarks

Openness

2 medium confidence
2.0
dashboard
public(free to browse)
eval-code
partial-open(several org repos MIT/Apache - SWE-bench docker registry, eci-public, MirrorCode-data)
benchmark-data
mixed(FrontierMath gated/private to prevent contamination
source
partial(public dashboard and MIT component repos, but no end-to-end harness published and the FrontierMath problems are unpublished)

Mixed: a public results dashboard plus some MIT- and Apache-licensed eval components, but the headline benchmark, FrontierMath, is deliberately private, and there is no single OSI-licensed end-to-end harness. That puts it closer to source-available than to fully open source.

  • https://github.com/epoch-research recorded 2026-06-04

    public org repos incl SWE-bench docker registry, eci-public (MIT), MirrorCode-data; mostly MIT/Apache-2.0

  • https://epoch.ai/benchmarks recorded 2026-06-04

    public benchmark dashboard running FrontierMath/GPQA Diamond/SWE-bench Verified/SimpleQA Verified across frontier models

  • https://epoch.ai/frontiermath recorded 2026-08-11

    Epoch describes FrontierMath Tiers 1-4 as 'a benchmark of several hundred unpublished, highly challenging mathematics problems'. The problems are held back by design, so the benchmark behind the public dashboard cannot be run by a reader.

  • https://api.github.com/orgs/epoch-research/repos recorded 2026-08-13

    33 public repos in the epoch-research org. Licensed eval components include SWE-bench and SWE-bench-c7d4d9d4 (MIT), eci-public (MIT), MirrorCode (MIT), benchmark-stitching (MIT), epochutils (MIT), ftm (MIT) and timelines-direct-approach (Apache-2.0); MirrorCode-data and many others carry no license at all. There is no single end-to-end harness repo among them.

  • https://epoch.ai/frontiermath recorded 2026-08-13

    Still 'A benchmark of several hundred unpublished, highly challenging mathematics problems', with Tiers 1-3 and Tier 4 deliberately withheld, so the headline benchmark behind the dashboard cannot be run by a reader.