Llamacha
unknownScores
1 product on the map — 1 open.
Llamacha monolingual Quechua corpus
Openness
4 medium confidence- license
- apache-2.0
- access
- public
- dataset_card
- partial(Card names the paper and license but its source, collection and curation sections are unfilled template text.)
The corpus downloads from Hugging Face without a gate under Apache-2.0. The card says nothing about where the non-Wikipedia text came from, which the QuBERT paper covers only in outline.
- https://huggingface.co/api/datasets/Llamacha/monolingual-quechua-iic recorded 2026-09-24
gated: false; license tag apache-2.0.
- https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic/raw/main/README.md recorded 2026-09-24
"license: apache-2.0"; "Licensing Information Apache-2.0"; Source Data sections read "More Information Needed".
Adoption
1 high confidenceHugging Face downloads over the trailing month for the one corpus repository.
- https://huggingface.co/api/datasets/Llamacha/monolingual-quechua-iic recorded 2026-09-24
482 downloads in the trailing 30 days for Llamacha/monolingual-quechua-iic
Capability
1 medium confidenceThe corpus gives Southern Quechua text gathered mainly in southern Peru plus the Quechua parts of Wikipedia and OSCAR, and its authors trained QuBERT and a Quechua RoBERTa on it, though no other model uses it and its card leaves sources largely undescribed. At under a hundred million tokens it is tiny beside English pretraining corpora of trillions of tokens, and a pooled multilingual corpus such as Glot's gathers more.
- https://aclanthology.org/2022.deeplo-1.1/ recorded 2026-09-24
"our curated corpus is created from text gathered from the southern region of Peru"; "a public, pre-trained, BERT model called QuBERT".
- https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic/raw/main/README.md recorded 2026-09-24
"Size of downloaded dataset files: 373.28 MB"; "This corpus also includes the Wiki and OSCAR corpora".