AI Potluck
Model components / Evaluation code

Artificial Analysis Intelligence Index

Artificial Analysis

Independent benchmarking service that runs a fixed suite of academic and agentic evaluations against every major model and publishes a composite Intelligence Index with confidence intervals. The composite weights four categories - agents, coding, scientific reasoning and general - unequally, and the methodology is published in full.

Several of the evaluations run on the Stirrup agent harness, which is released separately from the index pipeline. Verified 2026-08-13 via the Artificial Analysis benchmarking methodology page and the Stirrup license endpoint.

Openness

2 high confidence
2.0
platform
closed(proprietary benchmarking site + methodology)
source
partial(Stirrup covers one of ten evals
methodology
documented(public methodology page, per-eval breakdown, temp/token params)
datasets
closed(maintains internal copies of all eval datasets)
harness
partial-open(Stirrup MIT, used for the agentic GDPval-AA eval only, 413 stars)

The methodology is well documented and one component harness, Stirrup, is MIT-licensed, but the Index itself - pipeline, internal dataset copies, grader config - is proprietary. That is an open periphery rather than an open core: a repo exists while the runtime never ships, which makes the source only partial and the product source-available. The methodology page is at v4.1.1, with nine evals across four unequally weighted categories (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%), and Stirrup backs GDPval-AA v2, AA-Briefcase and Harvey LAB-AA rather than a single eval, so the open piece is somewhat larger than the components listed here suggest. Neither detail changes the judgment on source: Artificial Analysis still states 'We maintain internal copies of all evaluation datasets' and still publishes no end-to-end Index pipeline.

Adoption

4 medium confidence
4.0

The AA Intelligence Index is one of the most-cited independent model-ranking standards in frontier release coverage, referenced across vendor launches and the tech press as the go-to cross-model index, and this project's own capability scoring leans on it for records such as Grok 4.20 and Gemini 3.5 Flash. No unique-visitor count is published, so level 4 reflects its reach as a standard citation across the industry, banded conservatively and directionally.

Capability

4 high confidence
4.0

Frontier-grade composite coverage spanning agentic/coding/general/scientific with contamination-resistant internal datasets and standardized running. Strong capability. Not a 5: it is a curated index/methodology rather than an open community harness, and held-out/proprietary running limits external reproducibility.

Verified 2026-08-13