Llamacha monolingual Quechua corpus
LlamachaMonolingual-Quechua-IIC is a Southern Quechua text corpus for training language models, collected mainly from the southern region of Peru and combined with the Quechua portions of Wikipedia and OSCAR. Its authors used it to train QuBERT, a BERT model for Quechua, and tested that on named-entity recognition and part-of-speech tagging. The Llamacha group released it.
Openness
4 medium confidence- license
- apache-2.0
- access
- public
- dataset_card
- partial(Card names the paper and license but its source, collection and curation sections are unfilled template text.)
The corpus downloads from Hugging Face without a gate under Apache-2.0. The card says nothing about where the non-Wikipedia text came from, which the QuBERT paper covers only in outline.
- https://huggingface.co/api/datasets/Llamacha/monolingual-quechua-iic recorded 2026-09-24
gated: false; license tag apache-2.0.
- https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic/raw/main/README.md recorded 2026-09-24
"license: apache-2.0"; "Licensing Information Apache-2.0"; Source Data sections read "More Information Needed".
Adoption
1 high confidenceHugging Face downloads over the trailing month for the one corpus repository.
- https://huggingface.co/api/datasets/Llamacha/monolingual-quechua-iic recorded 2026-09-24
482 downloads in the trailing 30 days for Llamacha/monolingual-quechua-iic
Capability
1 medium confidenceThe corpus gives Southern Quechua text gathered mainly in southern Peru plus the Quechua parts of Wikipedia and OSCAR, and its authors trained QuBERT and a Quechua RoBERTa on it, though no other model uses it and its card leaves sources largely undescribed. At under a hundred million tokens it is tiny beside English pretraining corpora of trillions of tokens, and a pooled multilingual corpus such as Glot's gathers more.
- https://aclanthology.org/2022.deeplo-1.1/ recorded 2026-09-24
"our curated corpus is created from text gathered from the southern region of Peru"; "a public, pre-trained, BERT model called QuBERT".
- https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic/raw/main/README.md recorded 2026-09-24
"Size of downloaded dataset files: 373.28 MB"; "This corpus also includes the Wiki and OSCAR corpora".
Verified 2026-09-24