AI Potluck
Model components / Benchmark / eval datasets

OpenAI Internal Evals

OpenAI (Internal)

OpenAI's internal evaluation suites, used in model development and in safety assessment under its Preparedness Framework. They cover harder coding challenges, agentic task completion and cybersecurity-related tests, and are referenced in system cards and safety reports under the framework's tracked categories.

Distinct from the openai/evals harness, which is catalogued separately under evaluation_code. Verified 2026-08-13 via the OpenAI Preparedness Framework v2 PDF.

Openness

0 high confidence
0.0
data
closed(internal eval sets for frontier-risk Tracked Categories, never published)
datasheet
no
license
none
access
internal-only(only summarized/redacted results released alongside major deployments)

Fully internal/private eval datasets; only redacted result summaries are published, the underlying example sets are not redistributable.

  • https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf recorded 2026-08-13

    Preparedness Framework v2, 22 pages, "Last updated - 15th April, 2025". Tracked Categories are "those capabilities which we track most closely, measuring them during each covered deployment"; section 5.2 says "Such disclosures about results and safeguards may be redacted or summarized where necessary, such as to protect intellectual property or safety". No eval set, license or datasheet is published

Adoption

not assessed

Internal-only eval sets used solely by OpenAI for its own deployment decisions; no external users, no downloads, no third-party citation as a runnable standard. No honest external adoption signal exists.

Capability

not assessed

Capability is not a meaningful axis for an eval dataset, so it is left unscored. The published framework describes the process around the eval sets rather than any capability of the sets themselves.

Verified 2026-08-13