GAIA (General AI Assistants)
MetaBenchmark for general AI assistants from Meta AI and Hugging Face: 466 real-world questions across three difficulty levels requiring reasoning, multimodality, web browsing and tool use. Answers to 300 of them are withheld to power a submission-scored leaderboard, leaving a fully public development split.
Answers to the test split are held server-side, so a score cannot be reproduced locally. Verified 2026-08-13 via the GAIA dataset card on Hugging Face and the GAIA paper abstract.
Openness
2 high confidence- license
- not-clearly-stated-on-card(the Hub API cardData carries no license key and no license tag
- questions
- public(val)
- answers
- held-out(300-Q test,private)
- access
- HF-gated(accept-terms,anti-resharing)
Test answers are private and the dataset itself is access-gated. No license is stated anywhere the project publishes: the Hub API returns no license field and no license tag, the gated card states only the no-resharing condition, and the paper states no release terms. The license is therefore recorded as unstated rather than guessed at; either way, the held-out answers decide the score.
- https://huggingface.co/datasets/gaia-benchmark/GAIA recorded 2026-08-13
the page still says "You need to agree to share your contact information to access this dataset" and the extra_gated_prompt is the no-resharing condition; the card reads "Each level is divided into a fully public dev set for validation, and a test set with private answers and metadata"
- https://huggingface.co/api/datasets/gaia-benchmark/GAIA recorded 2026-08-13
cardData carries no license key and the tags array carries no license tag; gating is auto with a no-resharing condition
Adoption
3 high confidence17,975 Hugging Face downloads in the trailing 30 days for gaia-benchmark/GAIA, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models. GAIA is cited across agent evaluation harnesses, though its download volume has fallen sharply.
- https://huggingface.co/api/datasets/gaia-benchmark/GAIA recorded 2026-08-12
17,975 downloads in the trailing 30 days for gaia-benchmark/GAIA
Capability
5 high confidenceThe canonical end-to-end agentic-assistant benchmark: distinct from knowledge and code benchmarks, and contamination-resistant.
- https://arxiv.org/abs/2311.12983 recorded 2026-08-13
"GAIA - a benchmark for General AI Assistants"; the abstract still reads "we devise 466 questions and their answer. We release our questions while retaining answers to 300 of them to power a leader-board", and still reports "human respondents obtain 92% vs. 15% for GPT-4 equipped with plugins" over reasoning, multi-modality, web browsing and tool use
Verified 2026-08-12