GAIA (General AI Assistants)
MetaA benchmark for General AI Assistants from Meta AI / Hugging Face: 466 real-world questions across 3 difficulty levels requiring reasoning, multi-modality, web browsing, and tool use. Simple for humans (92%) but hard for frontier systems (GPT-4+plugins scored 15% at release). A 300-question test set with private answers powers a submission-based leaderboard.
Double-gated: test-set answers held out (server-side scoring) AND the HF dataset is access-gated (accept terms, anti-resharing clause). ~42k monthly downloads.
Openness
2 high confidence- questions
- public(val)
- answers
- held-out(300-Q test,private)
- access
- HF-gated(accept-terms,anti-resharing)
Test answers are private and the dataset itself is access-gated.
- https://huggingface.co/datasets/gaia-benchmark/GAIA recorded 2026-06-17
gated access, private test answers
Adoption
4 high confidence~42k monthly HF downloads; widely cited and integrated into agent eval harnesses.
- https://huggingface.co/datasets/gaia-benchmark/GAIA recorded 2026-06-17
~42,218 downloads/month
Capability
5 high confidenceThe canonical end-to-end agentic-assistant benchmark; distinct from knowledge and code benchmarks, contamination-resistant.
- https://arxiv.org/abs/2311.12983 recorded 2026-06-17
466 Q, 92% human vs 15% GPT-4, held-out answers
Unchanged since 2026-06-17 (last edited, not re-checked)