AI Potluck
Back to Gap Map Model components / Language-specific datasets

Jojajovai

PLN group, Universidad de la República
restricted / Overall score: 1.0

Jojajovai is a Guarani-Spanish parallel corpus of about 30,000 sentence pairs gathered from eight sources, among them news, blogs, books and the AmericasNLP data, each split into training, development and test sets. Native speakers annotated a test sample for Guarani dialect, including Jopara, and for translation correctness. Researchers at Universidad de la República in Uruguay built it with partners in Paraguay, Brazil and Spain.

The LREC 2022 paper the README cites was not opened, so the collection details come from the README.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(The repository has an MIT LICENSE file (Grupo PLN, UdelaR)
access
public
dataset_card
present

The corpus sits in a public GitHub repository with an MIT license file, but the README never says that license applies to the sentences, which come from several outside sources.

Adoption

1 low confidence
1.0

GitHub stars on the Jojajovai repository, the only public channel for the corpus; a star is attention rather than use.

Capability

1 medium confidence
1.0

Jojajovai aggregates Guarani-Spanish sentence pairs from eight sources, documented source by source and including the AmericasNLP Guarani data, with a native-speaker-checked test sample, though no named model or benchmark built on it turned up. At about thirty thousand sentence pairs it is minute beside the billions of pairs in the largest parallel collections, alongside the AmericasNLP shared-task data.

Verified 2026-09-24