OpenAI Internal Evals
OpenAI (Internal)OpenAI maintains proprietary internal evaluation suites for coding capability assessment, used in model development and safety evaluations under their Preparedness Framework. These include harder coding challenges, agentic task completions, and cybersecurity-related coding tests that are never publicly released. Referenced in system cards and safety reports but datasets remain private to prevent contamination and maintain evaluation integrity.
OpenAI internal dangerous-capability evaluations referenced in the Preparedness Framework (v2, updated 15 Apr 2025). The scored artifact is the internal eval sets used for Tracked Categories. Verified via the OpenAI Preparedness Framework v2 PDF (cdn.openai.com) and openai.com/index/updating-our-preparedness-framework. NOT the separate open source openai/evals harness (that is evaluation_code, not a dataset).
Openness
0 high confidence- data
- closed(internal eval sets for frontier-risk Tracked Categories, never published)
- datasheet
- no
- license
- none
- access
- internal-only(only summarized/redacted results released alongside major deployments)
Fully internal/private eval datasets; only redacted result summaries are published, the underlying example sets are not redistributable.
- https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf recorded 2026-06-04
Preparedness Framework v2 describes internal capability evaluations per Tracked Category with disclosures redacted/summarized to protect IP and safety
Adoption
not assessedInternal-only eval sets used solely by OpenAI for its own deployment decisions; no external users, no downloads, no third-party citation as a runnable standard. No honest external adoption signal exists.
- https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf recorded 2026-06-04
evaluations are internal to OpenAI's deployment process, not a shared/usable external dataset
Capability
not assessedCapability not a meaningful axis for an eval dataset.
Unchanged since 2026-06-09 (last edited, not re-checked)