Artificial Analysis
companyOpenness profile
1 product on the map — 1 open-ish.
Artificial Analysis Intelligence Index
Openness
2 high confidence- platform
- closed(proprietary benchmarking site + methodology)
- source
- partial(Stirrup covers one of ten evals
- methodology
- documented(public methodology page, per-eval breakdown, temp/token params)
- datasets
- closed(maintains internal copies of all eval datasets)
- harness
- partial-open(Stirrup MIT, used for the agentic GDPval-AA eval only, 413 stars)
The methodology is well documented and one component harness, Stirrup, is MIT-licensed, but the Index itself - pipeline, internal dataset copies, grader config - is proprietary. That is an open periphery rather than an open core: a repo exists while the runtime never ships, which makes the source only partial and the product source-available. The methodology page is at v4.1.1, with nine evals across four unequally weighted categories (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%), and Stirrup backs GDPval-AA v2, AA-Briefcase and Harvey LAB-AA rather than a single eval, so the open piece is somewhat larger than the components listed here suggest. Neither detail changes the judgment on source: Artificial Analysis still states 'We maintain internal copies of all evaluation datasets' and still publishes no end-to-end Index pipeline.
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-06-04
v4.0.4 (Mar 2026), 10-eval composite, internal copies of all datasets, standardized params (proprietary pipeline + open Stirrup harness for GDPval-AA)
- https://github.com/ArtificialAnalysis/Stirrup recorded 2026-06-04
Stirrup MIT-licensed agent harness, 413 stars (the only open component)
- https://artificialanalysis.ai/methodology/intelligence-benchmarking recorded 2026-08-13
Intelligence Index v4.1.1, nine evals across four unequally weighted categories - Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18% - namely GDPval-AA v2, tau-cubed Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity's Last Exam, GPQA Diamond and CritPt. 'We maintain internal copies of all evaluation datasets'. Stirrup is named as 'our open source agentic harness' for GDPval-AA v2, AA-Briefcase and Harvey LAB-AA. The Index pipeline itself is not published.
- https://api.github.com/repos/ArtificialAnalysis/Stirrup/license recorded 2026-08-13
GitHub's license endpoint reports spdx_id MIT for ArtificialAnalysis/Stirrup - the one open component - and returns the MIT text as the LICENSE body. 546 stars.