Artificial Analysis Intelligence Index
Artificial AnalysisIndependent benchmarking service that runs a fixed suite of academic and agentic evaluations against every major model and publishes a composite Intelligence Index with confidence intervals. The composite weights four categories - agents, coding, scientific reasoning and general - unequally, and the methodology is published in full.
Several of the evaluations run on the Stirrup agent harness, which is released separately from the index pipeline. Verified 2026-08-13 via the Artificial Analysis benchmarking methodology page and the Stirrup license endpoint.
Openness
2 high confidence- platform
- closed(proprietary benchmarking site + methodology)
- source
- partial(Stirrup covers one of ten evals
- methodology
- documented(public methodology page, per-eval breakdown, temp/token params)
- datasets
- closed(maintains internal copies of all eval datasets)
- harness
- partial-open(Stirrup MIT, used for the agentic GDPval-AA eval only, 413 stars)
The methodology is well documented and one component harness, Stirrup, is MIT-licensed, but the Index itself - pipeline, internal dataset copies, grader config - is proprietary. That is an open periphery rather than an open core: a repo exists while the runtime never ships, which makes the source only partial and the product source-available. The methodology page is at v4.1.1, with nine evals across four unequally weighted categories (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%), and Stirrup backs GDPval-AA v2, AA-Briefcase and Harvey LAB-AA rather than a single eval, so the open piece is somewhat larger than the components listed here suggest. Neither detail changes the judgment on source: Artificial Analysis still states 'We maintain internal copies of all evaluation datasets' and still publishes no end-to-end Index pipeline.
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-06-04
v4.0.4 (Mar 2026), 10-eval composite, internal copies of all datasets, standardized params (proprietary pipeline + open Stirrup harness for GDPval-AA)
- https://github.com/ArtificialAnalysis/Stirrup recorded 2026-06-04
Stirrup MIT-licensed agent harness, 413 stars (the only open component)
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-08-13
Intelligence Index v4.1.1, nine evals across four unequally weighted categories - Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18% - namely GDPval-AA v2, tau-cubed Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity's Last Exam, GPQA Diamond and CritPt. 'We maintain internal copies of all evaluation datasets'. Stirrup is named as 'our open source agentic harness' for GDPval-AA v2, AA-Briefcase and Harvey LAB-AA. The Index pipeline itself is not published.
- https://api.github.com/repos/ArtificialAnalysis/Stirrup/license recorded 2026-08-13
GitHub's license endpoint reports spdx_id MIT for ArtificialAnalysis/Stirrup - the one open component - and returns the MIT text as the LICENSE body. 546 stars.
Adoption
4 medium confidenceThe AA Intelligence Index is one of the most-cited independent model-ranking standards in frontier release coverage, referenced across vendor launches and the tech press as the go-to cross-model index, and this project's own capability scoring leans on it for records such as Grok 4.20 and Gemini 3.5 Flash. No unique-visitor count is published, so level 4 reflects its reach as a standard citation across the industry, banded conservatively and directionally.
- https://artificialanalysis.ai/ recorded 2026-06-04
live Intelligence Index leaderboard positioned as the cross-model standard reference
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-06-04
independent standardized index cited as the reference methodology
- https://artificialanalysis.ai/ recorded 2026-08-13
The live cross-model Intelligence Index leaderboard carries no unique-visitor, subscriber or usage count.
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-08-13
Methodology v4.1.1, published in full as the reference for the index. No adoption or traffic figure appears on it.
Capability
4 high confidenceFrontier-grade composite coverage spanning agentic/coding/general/scientific with contamination-resistant internal datasets and standardized running. Strong capability. Not a 5: it is a curated index/methodology rather than an open community harness, and held-out/proprietary running limits external reproducibility.
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-06-04
10-eval composite, 4 equal-weighted categories, standardized params, internal dataset copies
- https://artificialanalysis.ai/ recorded 2026-06-04
150+ models ranked on the Intelligence Index
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-08-13
Intelligence Index v4.1.1: nine evals - GDPval-AA v2, tau-cubed Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, HLE, GPQA Diamond, CritPt - across Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%. Standardized params and internal dataset copies; Stirrup is the open agentic harness behind GDPval-AA v2, AA-Briefcase and Harvey LAB-AA.
Verified 2026-08-13