AmericasNLP shared-task data
AmericasNLPThe AmericasNLP shared-task data pairs Spanish with ten Indigenous languages of the Americas, including Aymara, Bribri, Asháninka, Guarani, Wixarika, Nahuatl, Otomí, Quechua, Shipibo-Konibo and Rarámuri, for machine translation. Training sets gather existing parallel corpora, while development and test sets were translated by hand for the task. The AmericasNLP workshop organizers publish it with each shared task.
Openness
2 medium confidence- license
- not-clearly-stated-on-card(No LICENSE file in americasnlp2021 (LICENSE, LICENSE.md, LICENSE.txt all 404)
- access
- public
- dataset_card
- partial(README lists source corpora per language but does not describe the dev/test composition
The files sit in public GitHub repositories, but no license is stated, and the training portions come from corpora that each keep their own terms.
- https://aclanthology.org/2021.americasnlp-1.23/ recorded 2026-09-24
"We provided training sets consisting of data collected from various sources, as well as manually translated sentences for the development and test sets."
- https://raw.githubusercontent.com/AmericasNLP/americasnlp2021/main/README.md recorded 2026-09-24
"If you use one or more of the datasets included in this repository, please do not forget to cite each of te original papers." Lists sources per language; no license.
Adoption
1 low confidenceGitHub stars summed across the shared-task repositories, the only public channel for the data; a star is attention rather than use.
- https://ungh.cc/repos/AmericasNLP/americasnlp2021 recorded 2026-09-24
GitHub repository record for AmericasNLP/americasnlp2021 (via the ungh.cc mirror of the GitHub API): 45 stargazers, last push 2022-07-05.
- https://ungh.cc/repos/AmericasNLP/americasnlp2023 recorded 2026-09-24
GitHub repository record for AmericasNLP/americasnlp2023 (via the ungh.cc mirror of the GitHub API): 10 stargazers, last push 2023-05-15.
Capability
3 medium confidenceThe AmericasNLP data gives ten Indigenous languages of the Americas hand-translated test sets, documented in the shared-task papers, and the task's systems are scored on it. No use outside the shared tasks turned up and most training text is borrowed from other corpora, where FLORES+ underpins benchmarks and models well beyond one task.
- https://aclanthology.org/2021.americasnlp-1.23/ recorded 2026-09-24
"participants submitted machine translation systems for up to 10 indigenous languages. Overall, 8 teams participated with a total of 214 submissions."
- https://raw.githubusercontent.com/AmericasNLP/americasnlp2023/main/README.md recorded 2026-09-24
Baseline chrF table for aym, bzd, cni, gn, hch, nah, oto, quy, shp, tar; baseline is the Helsinki-NLP 2021 system.
Verified 2026-09-24