AI Potluck
Back to Gap Map Model components / Language-specific datasets

ChrEn

Shiyue Zhang
restricted / Overall score: 1.0

ChrEn is a Cherokee-English parallel corpus of about 14,000 sentence pairs for machine translation of an endangered language, with in-domain and out-of-domain splits and about 5,000 Cherokee monolingual sentences for semi-supervised training. Pairs come from published books and articles, and a Cherokee Old Testament set was added later. Shiyue Zhang, Benjamin Frey and Mohit Bansal at UNC released it.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(No LICENSE file (all variants 404)
access
public
dataset_card
present(Data README describes structure and splits

The files are in a public GitHub repository, but the text belongs to the original book and article authors, and the README offers it for research only, with copyright questions sent to a named contact.

Adoption

1 low confidence
1.0

GitHub stars on the ChrEn repository, the only public channel for the corpus; a star is attention rather than use.

Capability

1 medium confidence
1.0

ChrEn is the main open parallel corpus for Cherokee, aligning published translations with English and adding monolingual Cherokee sentences, and it is documented in an EMNLP paper, though no use beyond its authors' demo turned up. At about fourteen thousand sentence pairs it is minute beside the billions of pairs in the largest parallel collections, much like the AmericasNLP shared-task data.

Verified 2026-09-24