AI Potluck
Model components / Benchmark / eval datasets

OpenAI Internal Evals

OpenAI (Internal)

OpenAI maintains proprietary internal evaluation suites for coding capability assessment, used in model development and safety evaluations under their Preparedness Framework. These include harder coding challenges, agentic task completions, and cybersecurity-related coding tests that are never publicly released. Referenced in system cards and safety reports but datasets remain private to prevent contamination and maintain evaluation integrity.

OpenAI internal dangerous-capability evaluations referenced in the Preparedness Framework (v2, updated 15 Apr 2025). The scored artifact is the internal eval sets used for Tracked Categories. Verified via the OpenAI Preparedness Framework v2 PDF (cdn.openai.com) and openai.com/index/updating-our-preparedness-framework. NOT the separate open source openai/evals harness (that is evaluation_code, not a dataset).

Openness

0 high confidence
0.0
data
closed(internal eval sets for frontier-risk Tracked Categories, never published)
datasheet
no
license
none
access
internal-only(only summarized/redacted results released alongside major deployments)

Fully internal/private eval datasets; only redacted result summaries are published, the underlying example sets are not redistributable.

Adoption

not assessed

Internal-only eval sets used solely by OpenAI for its own deployment decisions; no external users, no downloads, no third-party citation as a runnable standard. No honest external adoption signal exists.

Capability

not assessed

Capability not a meaningful axis for an eval dataset.

Unchanged since 2026-06-09 (last edited, not re-checked)