Introducing the Gap Map v0.1
To create the map, we used both a discovery step (to find the universe) and a more rigorous scoring and enrichment step (to grade each product). Learn more about our methodology.
We track over 24,600 open source AI artifacts across the stack. This map scores more than 421 of them in depth on three independent axes:
- Openness (graded 0–5 against openness frameworks, not a yes/no: the Model Openness Framework for models, OSI classes for software, with data and hardware analogues)
- Adoption (real usage, not stars)
- Capability (benchmarks where they exist, feature coverage where they don’t)
Every score is sourced. The remaining unscored products are the uncategorized long tail, tracked by usage signal but not yet scored. The openness framework descends directly from the 2024 Columbia Convening on Openness in AI.
Introduction
To create the map, we used both a discovery step (to find the universe) and a more rigorous scoring and enrichment step (to grade each product). The taxonomy we use to categorize products descends directly from the 2024 Columbia Convening on Openness in AI.
General-purpose AI system stack & Dimensions of Openness from the 2024 Columbia Convening on Openness in AI.
The framework has two levels of analysis:
- Each product is scored on three axes: openness, adoption, and capability.
- Each category is rolled up from its products' scores into a maturity stage from 0 (Void) to 5 (Mature), plus a set of gaps naming what its open ecosystem still lacks.
Notebooks
Besides this gap map, the underlying data is also available to explore directly through our notebooks. Each notebook lets you explore a specific slice of the dataset. You can find the notebooks here:
All notebooks are also available on GitHub at currentai-org/os-ai-map, where you can inspect the underlying code, or adapt them for your own analysis.
Discovery and scoring
The discovery step identifies the universe of candidate products and artifacts; the scoring step enriches and grades a curated subset of them.
The discovery step draws on large-scale open data from the software supply chain compiled by Open Source Observer. We seeded it from Chip Huyen's Good AI List, a catalog of AI-focused repositories, then broadened it through analysis of the Hugging Face Hub, the Open LLM Leaderboard, the AI Incident Database, package registries and SBOMs, and academic and industry publications. From these sources we assembled approximately 24,821 candidate products, including 15,375 GitHub repositories, 6,428 models and datasets, and 2,823 package entries. We ranked the candidates by adoption signal (repository stars, package and model downloads, and related measures) and enriched the most prominent first.
The scoring step enriched and graded 421 products in depth: 266 software tools and libraries, 85 models, 50 datasets, and 20 hardware projects, produced by 228 organizations. We organize these products into 14 categories across 3 layers of the stack (model components, product / UX, and infrastructure), though we do not cover the stack exhaustively. The remaining 24,400 artifacts constitute the uncategorized long tail: they are tracked by usage signal but carry no openness, adoption, or capability score until they are researched and cited.
This is v1 of the Gap Map, to contribute, request a correction, or suggest an improvement, find us on Git Hub.