AI Potluck
Model components / Training & synthetic datasets

Aya Collection

Cohere

The Aya Collection is a large multilingual instruction-tuning dataset spanning 114+ languages, built from templated and translated subsets for tasks like classification, summarization, and translation. It was created by Cohere Labs (formerly Cohere For AI) as part of the Aya open-science initiative. The collection holds hundreds of millions of examples across 64 subsets.

Verified live 2026-06-22 via primary sources. Ungated, Apache-2.0 licensed, with a complete dataset card.

Openness

5 high confidence
5.0
license
apache-2.0
gated
false
card
present

Ungated, Apache-2.0 licensed, with a complete dataset card.

Adoption

3 high confidence
3.0

26,509 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 high confidence
4.0

Instruction collection behind Aya 101 (2024), Aya 23, and Aya Expanse, which beat Gemma 2 / Qwen 2.5 / Llama 3.1 in their size class; strong, multiple notable open models.

Unchanged since 2026-07-04 (last edited, not re-checked)