AI Potluck
Model components / Evaluation code

OpenAI Evals API & Dashboard

OpenAI

Managed evaluation product inside the OpenAI Platform, Evals API plus a Dashboard UI that lets developers define eval datasets, run them against any OpenAI model, grade with custom or built-in graders, and track regressions across versions. Picked by teams already on OpenAI who want the 'measure → improve → ship' loop without standing up infra. Pricing rolls into standard OpenAI API token usage.

Hosted, proprietary evaluation product on the OpenAI platform (API + web dashboard), distinct from the open source openai/evals repo. Lets developers define grading criteria, upload JSONL test sets, run against Chat Completions/Responses APIs, and inspect pass/fail + cost metrics. BEING SUNSET: dashboard read-only 31 Oct 2026, full shutdown 30 Nov 2026 (OpenAI recommends Datasets as replacement). Confirmed live June 2026 via OpenAI developer docs.

Openness

1 high confidence
1.0
license
Proprietary(SaaS, OpenAI-platform-only)
source
closed

Closed hosted SaaS bound to the OpenAI platform; not the MIT-licensed openai/evals repo (that is scored separately).

Adoption

not assessed

OpenAI publishes no usage figures for the Evals API/dashboard specifically, and the product is being sunset, so there is no honest adoption signal to report.

Capability

2 medium confidence
2.0

Convenient managed eval surface for OpenAI-model users, but narrow (single-provider, no benchmark library) and on a deprecation path, not a frontier harness.

Unchanged since 2026-06-24 (last edited, not re-checked)