OpenAI Evals API & Dashboard
OpenAIManaged evaluation product inside the OpenAI platform: an Evals API and a dashboard that let developers define eval datasets, run them against OpenAI models, grade with built-in or custom graders, and inspect pass/fail and cost metrics across versions.
Being sunset: read-only from 31 October 2026 and shut down on 30 November 2026, with OpenAI directing users to Datasets. Distinct from the openai/evals repository, which is a separate record. Verified 2026-08-13 via the OpenAI Evals guide.
Openness
1 high confidence- license
- Proprietary(SaaS, OpenAI-platform-only)
- source
- closed
Closed hosted SaaS bound to the OpenAI platform; not the MIT-licensed openai/evals repo (that is scored separately).
- https://developers.openai.com/api/docs/guides/evals recorded 2026-06-04
hosted proprietary Evals API + dashboard on OpenAI platform; sunset dates (read-only Oct 31 2026, shutdown Nov 30 2026)
- https://developers.openai.com/api/docs/guides/evals recorded 2026-08-13
The guide states 'OpenAI is deprecating the Evals platform' and 'Evals will become read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026', and points new users at Datasets instead. It is documented purely as a hosted surface of the OpenAI API - no repository, no self-hostable implementation and no license to any evaluation code is offered anywhere on the page.
Adoption
not assessedOpenAI publishes no usage figures for the Evals API/dashboard specifically, and the product is being sunset, so there is no honest adoption signal to report.
- https://developers.openai.com/api/docs/guides/evals recorded 2026-06-04
no usage/customer figures; product on sunset path
- https://developers.openai.com/api/docs/guides/evals recorded 2026-08-13
The guide publishes no eval-run count, no customer count and no developer count for the Evals platform, and the deprecation notice adds none. Nothing bandable appears on it.
Capability
2 medium confidenceA convenient managed eval surface for people already using OpenAI models, but narrow - single provider, no benchmark library - and on a published deprecation path rather than a frontier harness.
- https://developers.openai.com/api/docs/guides/evals recorded 2026-06-04
grading methods, dataset upload, run execution vs OpenAI APIs, dashboard + cost analysis
- https://developers.openai.com/api/docs/guides/evals recorded 2026-08-13
Criteria-based grading against OpenAI models only, dataset upload, runs and a dashboard; no multi-provider backend and no standardized benchmark suite. Read-only from 2026-10-31, shut down 2026-11-30.
Verified 2026-08-13