AI Potluck
Back to Gap Map Model components / Language-specific datasets

Llamacha monolingual Quechua corpus

Llamacha
open / Overall score: 1.0

Monolingual-Quechua-IIC is a Southern Quechua text corpus for training language models, collected mainly from the southern region of Peru and combined with the Quechua portions of Wikipedia and OSCAR. Its authors used it to train QuBERT, a BERT model for Quechua, and tested that on named-entity recognition and part-of-speech tagging. The Llamacha group released it.

Openness

4 medium confidence
4.0
license
apache-2.0
access
public
dataset_card
partial(Card names the paper and license but its source, collection and curation sections are unfilled template text.)

The corpus downloads from Hugging Face without a gate under Apache-2.0. The card says nothing about where the non-Wikipedia text came from, which the QuBERT paper covers only in outline.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing month for the one corpus repository.

Capability

1 medium confidence
1.0

The corpus gives Southern Quechua text gathered mainly in southern Peru plus the Quechua parts of Wikipedia and OSCAR, and its authors trained QuBERT and a Quechua RoBERTa on it, though no other model uses it and its card leaves sources largely undescribed. At under a hundred million tokens it is tiny beside English pretraining corpora of trillions of tokens, and a pooled multilingual corpus such as Glot's gathers more.

Verified 2026-09-24