Jojajovai
PLN group, Universidad de la RepúblicaJojajovai is a Guarani-Spanish parallel corpus of about 30,000 sentence pairs gathered from eight sources, among them news, blogs, books and the AmericasNLP data, each split into training, development and test sets. Native speakers annotated a test sample for Guarani dialect, including Jopara, and for translation correctness. Researchers at Universidad de la República in Uruguay built it with partners in Paraguay, Brazil and Spain.
The LREC 2022 paper the README cites was not opened, so the collection details come from the README.
Openness
2 medium confidence- license
- not-clearly-stated-on-card(The repository has an MIT LICENSE file (Grupo PLN, UdelaR)
- access
- public
- dataset_card
- present
The corpus sits in a public GitHub repository with an MIT license file, but the README never says that license applies to the sentences, which come from several outside sources.
- https://raw.githubusercontent.com/pln-fing-udelar/jojajovai/main/LICENSE recorded 2026-09-24
MIT License, "Copyright (c) 2022 Grupo PLN, UdelaR".
- https://raw.githubusercontent.com/pln-fing-udelar/jojajovai/main/README.md recorded 2026-09-24
"a Guarani-Spanish parallel corpus of about 30,000 sentence pairs, structured as a set of different sources"; no license statement; "The file jojajovai_all.csv contains the data".
Adoption
1 low confidenceGitHub stars on the Jojajovai repository, the only public channel for the corpus; a star is attention rather than use.
- https://ungh.cc/repos/pln-fing-udelar/jojajovai recorded 2026-09-24
GitHub repository record for pln-fing-udelar/jojajovai (via the ungh.cc mirror of the GitHub API): 21 stargazers, last push 2022-07-05.
Capability
1 medium confidenceJojajovai aggregates Guarani-Spanish sentence pairs from eight sources, documented source by source and including the AmericasNLP Guarani data, with a native-speaker-checked test sample, though no named model or benchmark built on it turned up. At about thirty thousand sentence pairs it is minute beside the billions of pairs in the largest parallel collections, alongside the AmericasNLP shared-task data.
- https://raw.githubusercontent.com/pln-fing-udelar/jojajovai/main/README.md recorded 2026-09-24
Source table totals 30,855 pairs: 20,207 train, 5,314 dev, 5,334 test; "Three native annotators were given a sample of sentence pairs from each set".
Verified 2026-09-24