AI Potluck
Model components / Evaluation code

Inspect AI

UK AI Security Institute

Frontier-AI evaluation framework with one interface over OpenAI, Anthropic, Google, xAI, Bedrock and local vLLM/Ollama backends, plus a collection of 200+ pre-built evals (Inspect Evals) co-maintained with Arcadia Impact and the Vector Institute. Picked by AISI, EU AI Office, and many frontier labs as the standard for safety/dangerous-capability evaluations. 2.1k stars but extremely active, 1,333 commits in last 90 days, 227 contributors.

UKGovernmentBEIS/inspect_ai (UK AI Security Institute + Meridian Labs), MIT, ~2.2k stars, ~5,886 commits, very actively developed. Framework for frontier LLM/agent evaluations: 200+ pre-built evals, tool use, multi-turn, model-graded scoring, one interface over OpenAI/Anthropic/Google/xAI/Bedrock/Azure/local vLLM/Ollama. Confirmed live June 2026.

Openness

5 high confidence
5.0
license
MIT(OSI)
source
public(full framework, 200+ evals)
governance
public-sector(UK AISI)+Meridian Labs
core-gated
ungated

Fully OSI (MIT), government-backed, no proprietary core, full pipeline open.

Adoption

4 high confidence
4.0

Headline signal (per recipe: adoption by frontier labs / cited in release-grade safety work): adopted as the evaluation framework of choice by major labs incl. Anthropic, Google DeepMind, and xAI/Grok, and powers nearly all of UK AISI's automated frontier-model evaluations. Corroborated by ~9.9M PyPI downloads in the last 30 days (pepy.tech) / 10.5M (pypistats), though that figure is inflated by CI/dependency traffic, so lab adoption is treated as the primary basis. Placed at level 4, not 5, because the human-operator base (eval engineers at labs/AISI) is far smaller than mass-market scale.

Capability

5 high confidence
5.0

Frontier-defining for this category: combines the broadest backend coverage, agentic-eval support, and de-facto-standard status among frontier-safety labs. One of the 1-2 leaders, so scored 5.

Unchanged since 2026-07-30 (last edited, not re-checked)