Inspect AI
UK AI Security InstituteFrontier-AI evaluation framework with one interface over OpenAI, Anthropic, Google, xAI, Bedrock and local vLLM/Ollama backends, plus a collection of 200+ pre-built evals (Inspect Evals) co-maintained with Arcadia Impact and the Vector Institute. Picked by AISI, EU AI Office, and many frontier labs as the standard for safety/dangerous-capability evaluations. 2.1k stars but extremely active, 1,333 commits in last 90 days, 227 contributors.
UKGovernmentBEIS/inspect_ai (UK AI Security Institute + Meridian Labs), MIT, ~2.2k stars, ~5,886 commits, very actively developed. Framework for frontier LLM/agent evaluations: 200+ pre-built evals, tool use, multi-turn, model-graded scoring, one interface over OpenAI/Anthropic/Google/xAI/Bedrock/Azure/local vLLM/Ollama. Confirmed live June 2026.
Openness
5 high confidence- license
- MIT(OSI)
- source
- public(full framework, 200+ evals)
- governance
- public-sector(UK AISI)+Meridian Labs
- core-gated
- ungated
Fully OSI (MIT), government-backed, no proprietary core, full pipeline open.
- https://github.com/UKGovernmentBEIS/inspect_ai recorded 2026-06-04
MIT license; framework for LLM evals with 200+ pre-built evals; UK AISI maintainer; 2.2k stars, 5,886 commits
Adoption
4 high confidenceHeadline signal (per recipe: adoption by frontier labs / cited in release-grade safety work): adopted as the evaluation framework of choice by major labs incl. Anthropic, Google DeepMind, and xAI/Grok, and powers nearly all of UK AISI's automated frontier-model evaluations. Corroborated by ~9.9M PyPI downloads in the last 30 days (pepy.tech) / 10.5M (pypistats), though that figure is inflated by CI/dependency traffic, so lab adoption is treated as the primary basis. Placed at level 4, not 5, because the human-operator base (eval engineers at labs/AISI) is far smaller than mass-market scale.
- https://www.aisi.gov.uk/blog/our-first-year recorded 2026-06-04
AISI runs thousands of frontier-model evaluations on Inspect; broad lab adoption
- https://inspect.aisi.org.uk/ recorded 2026-06-04
framework for frontier AI evaluations by UK AISI + Meridian Labs; adopted by major labs
- https://pepy.tech/projects/inspect-ai recorded 2026-06-04
87.89M total / 9,856,617 last-30-day PyPI downloads (corroboration)
Capability
5 high confidenceFrontier-defining for this category: combines the broadest backend coverage, agentic-eval support, and de-facto-standard status among frontier-safety labs. One of the 1-2 leaders, so scored 5.
- https://inspect.aisi.org.uk/models.html recorded 2026-06-04
single interface over OpenAI/Anthropic/Google/xAI/Bedrock/Azure/local vLLM/Ollama/llama-cpp
- https://github.com/UKGovernmentBEIS/inspect_ai recorded 2026-06-04
200+ pre-built evals; tool use, multi-turn, model-graded eval components
Unchanged since 2026-07-30 (last edited, not re-checked)