AI Potluck
Model components / Benchmark / eval datasets

GAIA (General AI Assistants)

Meta

A benchmark for General AI Assistants from Meta AI / Hugging Face: 466 real-world questions across 3 difficulty levels requiring reasoning, multi-modality, web browsing, and tool use. Simple for humans (92%) but hard for frontier systems (GPT-4+plugins scored 15% at release). A 300-question test set with private answers powers a submission-based leaderboard.

Double-gated: test-set answers held out (server-side scoring) AND the HF dataset is access-gated (accept terms, anti-resharing clause). ~42k monthly downloads.

Openness

2 high confidence
2.0
questions
public(val)
answers
held-out(300-Q test,private)
access
HF-gated(accept-terms,anti-resharing)

Test answers are private and the dataset itself is access-gated.

Adoption

4 high confidence
4.0

~42k monthly HF downloads; widely cited and integrated into agent eval harnesses.

Capability

5 high confidence
5.0

The canonical end-to-end agentic-assistant benchmark; distinct from knowledge and code benchmarks, contamination-resistant.

Unchanged since 2026-06-17 (last edited, not re-checked)