Aya Collection
CohereThe Aya Collection is a large multilingual instruction-tuning dataset spanning 114+ languages, built from templated and translated subsets for tasks like classification, summarization, and translation. It was created by Cohere Labs (formerly Cohere For AI) as part of the Aya open-science initiative. The collection holds hundreds of millions of examples across 64 subsets.
Verified live 2026-06-22 via primary sources. Ungated, Apache-2.0 licensed, with a complete dataset card.
Openness
5 high confidence- license
- apache-2.0
- gated
- false
- card
- present
Ungated, Apache-2.0 licensed, with a complete dataset card.
- https://huggingface.co/datasets/CohereLabs/aya_collection recorded 2026-06-22
Card describing 114+ languages, templated/translated subsets, Apache-2.0, ungated
Adoption
3 high confidence26,509 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/CohereLabs/aya_collection recorded 2026-07-04
downloads field = 26,509 (30-day)
Capability
4 high confidenceInstruction collection behind Aya 101 (2024), Aya 23, and Aya Expanse, which beat Gemma 2 / Qwen 2.5 / Llama 3.1 in their size class; strong, multiple notable open models.
- https://huggingface.co/blog/aya-expanse recorded 2026-07-04
Aya Expanse 8B/32B outperform Gemma 2, Qwen 2.5, Llama 3.1 in class (up to 76.6% win-rate on Arena-Hard-Auto)
- https://ritvik19.medium.com/papers-explained-151-aya-23-d01605c3ee80 recorded 2026-07-04
Aya Collection (513M examples, 114+ langs) feeds Aya 101 and Aya 23/Expanse instruction data
Unchanged since 2026-07-04 (last edited, not re-checked)