AI Potluck
Back to Gap Map Model components / Language-specific datasets

AmericasNLP shared-task data

AmericasNLP
restricted / Overall score: 2.4

The AmericasNLP shared-task data pairs Spanish with ten Indigenous languages of the Americas, including Aymara, Bribri, Asháninka, Guarani, Wixarika, Nahuatl, Otomí, Quechua, Shipibo-Konibo and Rarámuri, for machine translation. Training sets gather existing parallel corpora, while development and test sets were translated by hand for the task. The AmericasNLP workshop organizers publish it with each shared task.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(No LICENSE file in americasnlp2021 (LICENSE, LICENSE.md, LICENSE.txt all 404)
access
public
dataset_card
partial(README lists source corpora per language but does not describe the dev/test composition

The files sit in public GitHub repositories, but no license is stated, and the training portions come from corpora that each keep their own terms.

Adoption

1 low confidence
1.0

GitHub stars summed across the shared-task repositories, the only public channel for the data; a star is attention rather than use.

Capability

3 medium confidence
3.0

The AmericasNLP data gives ten Indigenous languages of the Americas hand-translated test sets, documented in the shared-task papers, and the task's systems are scored on it. No use outside the shared tasks turned up and most training text is borrowed from other corpora, where FLORES+ underpins benchmarks and models well beyond one task.

Verified 2026-09-24