AI Potluck
Model components / Evaluation code

OpenAI Evals API & Dashboard

OpenAI

Managed evaluation product inside the OpenAI platform: an Evals API and a dashboard that let developers define eval datasets, run them against OpenAI models, grade with built-in or custom graders, and inspect pass/fail and cost metrics across versions.

Being sunset: read-only from 31 October 2026 and shut down on 30 November 2026, with OpenAI directing users to Datasets. Distinct from the openai/evals repository, which is a separate record. Verified 2026-08-13 via the OpenAI Evals guide.

Openness

1 high confidence
1.0
license
Proprietary(SaaS, OpenAI-platform-only)
source
closed

Closed hosted SaaS bound to the OpenAI platform; not the MIT-licensed openai/evals repo (that is scored separately).

  • https://developers.openai.com/api/docs/guides/evals recorded 2026-06-04

    hosted proprietary Evals API + dashboard on OpenAI platform; sunset dates (read-only Oct 31 2026, shutdown Nov 30 2026)

  • https://developers.openai.com/api/docs/guides/evals recorded 2026-08-13

    The guide states 'OpenAI is deprecating the Evals platform' and 'Evals will become read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026', and points new users at Datasets instead. It is documented purely as a hosted surface of the OpenAI API - no repository, no self-hostable implementation and no license to any evaluation code is offered anywhere on the page.

Adoption

not assessed

OpenAI publishes no usage figures for the Evals API/dashboard specifically, and the product is being sunset, so there is no honest adoption signal to report.

Capability

2 medium confidence
2.0

A convenient managed eval surface for people already using OpenAI models, but narrow - single provider, no benchmark library - and on a published deprecation path rather than a frontier harness.

Verified 2026-08-13