ChrEn
Shiyue ZhangChrEn is a Cherokee-English parallel corpus of about 14,000 sentence pairs for machine translation of an endangered language, with in-domain and out-of-domain splits and about 5,000 Cherokee monolingual sentences for semi-supervised training. Pairs come from published books and articles, and a Cherokee Old Testament set was added later. Shiyue Zhang, Benjamin Frey and Mohit Bansal at UNC released it.
Openness
2 medium confidence- license
- not-clearly-stated-on-card(No LICENSE file (all variants 404)
- access
- public
- dataset_card
- present(Data README describes structure and splits
The files are in a public GitHub repository, but the text belongs to the original book and article authors, and the README offers it for research only, with copyright questions sent to a named contact.
- https://raw.githubusercontent.com/ZhangShiyue/ChrEn/main/data/README.md recorded 2026-09-24
"The copyright of the data belongs to original book/article authors or translators (hence, used for research purpose; and please contact Dr. Benjamin Frey for other copyright questions)."
- https://raw.githubusercontent.com/ZhangShiyue/ChrEn/main/README.md recorded 2026-09-24
"This repository contains the data/code/demo for the following paper"; links data, code and demo.
Adoption
1 low confidenceGitHub stars on the ChrEn repository, the only public channel for the corpus; a star is attention rather than use.
- https://ungh.cc/repos/ZhangShiyue/ChrEn recorded 2026-09-24
GitHub repository record for ZhangShiyue/ChrEn (via the ungh.cc mirror of the GitHub API): 24 stargazers, last push 2022-01-18.
Capability
1 medium confidenceChrEn is the main open parallel corpus for Cherokee, aligning published translations with English and adding monolingual Cherokee sentences, and it is documented in an EMNLP paper, though no use beyond its authors' demo turned up. At about fourteen thousand sentence pairs it is minute beside the billions of pairs in the largest parallel collections, much like the AmericasNLP shared-task data.
- https://arxiv.org/abs/2010.04791 recorded 2026-09-24
"ChrEn is extremely low-resource, only containing 14k sentence pairs in total"; "We also collect 5k Cherokee monolingual data".
- https://raw.githubusercontent.com/ZhangShiyue/ChrEn/main/data/README.md recorded 2026-09-24
Lists parallel splits, monolingual data, a Cherokee Old Testament addition and demo data.
Verified 2026-09-24