AI Potluck
Back to Gap Map Organization

Institute for Research and Innovation in Intelligent Systems (IRIIS)

lab

Scores

1 product on the map — 1 open.

Nepali Text Corpus

Openness

5 medium confidence
5.0
license
mit(Hugging Face card metadata only
access
public(Hugging Face gated: false)
dataset_card
present(brief card naming sources, article count and size)

The corpus downloads without a gate, and its card metadata tags it MIT, a software license, with no further terms for the scraped news and blog text. The card itself is a short summary, with collection details left to the paper.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single corpus repository. A download does not show whether the data was used to train a released model.

Capability

2 medium confidence
2.0

The corpus gives Nepali a crawled article collection larger than earlier Nepali corpora, described in a paper, and IRIIS Research trained its Nepali BERT, RoBERTa and GPT models on it, though no outside model is known to use it. At several billion tokens, estimated from its size on disk, it is far below the trillions of tokens in English pretraining corpora, and Sangraha, a documented corpus across twenty-two Indic languages, is far larger.