AI Potluck
Back to Gap Map Organization

Llamacha

unknown

Scores

1 product on the map — 1 open.

Llamacha monolingual Quechua corpus

Openness

4 medium confidence
4.0
license
apache-2.0
access
public
dataset_card
partial(Card names the paper and license but its source, collection and curation sections are unfilled template text.)

The corpus downloads from Hugging Face without a gate under Apache-2.0. The card says nothing about where the non-Wikipedia text came from, which the QuBERT paper covers only in outline.

Adoption

1 high confidence
1.0

Hugging Face downloads over the trailing month for the one corpus repository.

Capability

1 medium confidence
1.0

The corpus gives Southern Quechua text gathered mainly in southern Peru plus the Quechua parts of Wikipedia and OSCAR, and its authors trained QuBERT and a Quechua RoBERTa on it, though no other model uses it and its card leaves sources largely undescribed. At under a hundred million tokens it is tiny beside English pretraining corpora of trillions of tokens, and a pooled multilingual corpus such as Glot's gathers more.