AI Potluck
Model components / Benchmark / eval datasets

Scale SEAL Leaderboard

Scale AI

Scale AI's private evaluation sets behind its SEAL leaderboards, covering domains including instruction-following, math, coding and multilinguality. Problems are held private to prevent contamination and are scored by expert annotators, and the leaderboards rank frontier models rather than publishing the prompts.

Launched in 2024 and extended with new domains since. Verified 2026-08-13 via Scale AI's leaderboard announcement.

Openness

1 high confidence
1.0
data
closed(proprietary held-out prompt sets, ~1,000 examples/domain, deliberately unpublished to prevent contamination/training-set leakage)
datasheet
no(methodology described but examples not redistributable)
license
none(not distributed)
access
held-out(results visible on leaderboard, dataset not downloadable)

The eval sets are intentionally private and held out, which is the closed end of the data-openness scale: only the leaderboard results are published, never the data behind them. Two details recorded here do not appear anywhere on the page - the roughly 1,000 examples per domain, and the mixing in of some open-source sets - so the 1,000-example figure is not backed by the source cited for it.

  • https://scale.com/blog/leaderboard recorded 2026-08-13

    "Leaderboard Integrity - Unlike most public benchmarks, Scale's proprietary datasets will remain private and unpublished, ensuring they cannot be exploited or incorporated into model training data." No download link, no license statement and no dataset card anywhere on the page

Adoption

3 medium confidence
3.0

SEAL is a widely referenced expert-driven private leaderboard, cited across the LLM evaluation community as a contamination-resistant ranking, with broad press and industry attention. There is no published download or usage count, because the data is held out by design, so the band rests on its standing as an industry reference rather than on a measured usage figure.

  • https://scale.com/blog/leaderboard recorded 2026-08-13

    SEAL positioned as a trusted frontier-model leaderboard answering "Contamination & Overfitting" and "Inconsistent Reporting" in model comparison; no download, user or submission count is published anywhere on the page

Capability

not assessed

Capability is not a meaningful axis for an eval dataset, so it is left unscored; openness and adoption carry this category.

  • https://scale.com/blog/leaderboard recorded 2026-08-13

    the scored artifact is the private held-out prompt sets behind the leaderboard; the page ranks models rather than describing any capability of the prompt sets themselves

Verified 2026-08-13