OpenAI Internal Evals
OpenAI (Internal)OpenAI's internal evaluation suites, used in model development and in safety assessment under its Preparedness Framework. They cover harder coding challenges, agentic task completion and cybersecurity-related tests, and are referenced in system cards and safety reports under the framework's tracked categories.
Distinct from the openai/evals harness, which is catalogued separately under evaluation_code. Verified 2026-08-13 via the OpenAI Preparedness Framework v2 PDF.
Openness
0 high confidence- data
- closed(internal eval sets for frontier-risk Tracked Categories, never published)
- datasheet
- no
- license
- none
- access
- internal-only(only summarized/redacted results released alongside major deployments)
Fully internal/private eval datasets; only redacted result summaries are published, the underlying example sets are not redistributable.
- https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf recorded 2026-08-13
Preparedness Framework v2, 22 pages, "Last updated - 15th April, 2025". Tracked Categories are "those capabilities which we track most closely, measuring them during each covered deployment"; section 5.2 says "Such disclosures about results and safeguards may be redacted or summarized where necessary, such as to protect intellectual property or safety". No eval set, license or datasheet is published
Adoption
not assessedInternal-only eval sets used solely by OpenAI for its own deployment decisions; no external users, no downloads, no third-party citation as a runnable standard. No honest external adoption signal exists.
- https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf recorded 2026-08-13
evaluations are run internally as part of OpenAI's own deployment decisions, with no released artifact and no usage figure of any kind, so there is nothing external to band
Capability
not assessedCapability is not a meaningful axis for an eval dataset, so it is left unscored. The published framework describes the process around the eval sets rather than any capability of the sets themselves.
- https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf recorded 2026-08-13
the document describes thresholds and safeguards for model capability, never a capability of the eval sets themselves, and publishes no feature matrix or benchmark row for them
Verified 2026-08-13