INCLUDE
CohereINCLUDE is a multilingual knowledge and reasoning benchmark of four-option multiple-choice questions taken from local academic, professional and licensing exams in 44 languages. Because the questions come from each region's own exams rather than translated English tests, many probe regional knowledge. A smaller lite subset covers the same languages. It was built by EPFL and Cohere Labs researchers.
Openness
5 high confidence- license
- apache-2.0
- access
- public
- dataset_card
- present
Both repositories download from Hugging Face without a gate under Apache-2.0, and the full question set including answers is published.
- https://huggingface.co/api/datasets/CohereLabs/include-base-44 recorded 2026-09-24
gated: false; license tag apache-2.0.
- https://huggingface.co/api/datasets/CohereLabs/include-lite-44 recorded 2026-09-24
gated: false; license tag apache-2.0.
- https://huggingface.co/datasets/CohereLabs/include-base-44/raw/main/README.md recorded 2026-09-24
"It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics"; schema includes "answer".
Adoption
3 high confidenceHugging Face downloads over the trailing month for the base and lite repositories combined. Runs through evaluation harnesses that cache the data are counted only when they download it.
- https://huggingface.co/api/datasets/CohereLabs/include-base-44 recorded 2026-09-24
12370 downloads in the trailing 30 days for CohereLabs/include-base-44
- https://huggingface.co/api/datasets/CohereLabs/include-lite-44 recorded 2026-09-24
2614 downloads in the trailing 30 days for CohereLabs/include-lite-44
Capability
4 medium confidenceINCLUDE draws its questions from each region's own exams rather than translated English tests, is documented in its paper, and ships as a task in EleutherAI's evaluation harness. Its 44 languages are fewer than the 122 variants of Belebele.
- https://arxiv.org/abs/2411.19799 recorded 2026-09-24
"we construct an evaluation suite of 197,243 QA pairs from local exam sources".
- https://huggingface.co/datasets/CohereLabs/include-base-44/raw/main/README.md recorded 2026-09-24
Model performance table: Llama3.1-70B-Instruct 70.6, Qwen2.5-14B 62.3, Aya-expanse-32b 59.1.
- https://raw.githubusercontent.com/EleutherAI/lm-evaluation-harness/main/lm_eval/tasks/include/README.md recorded 2026-09-24
"INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages"; points to CohereForAI/include-base-44.
Verified 2026-09-24