AI Potluck
Back to Gap Map Model components / Language-specific datasets

Nepali Text Corpus

Institute for Research and Innovation in Intelligent Systems (IRIIS)
open / Overall score: 1.7

The Nepali Text Corpus is a monolingual Nepali pretraining corpus of about 6.4 million articles, 27.5 GB as CSV, collected from Nepali news websites, blogs and other online sources. IRIIS Research assembled it to pretrain Nepali BERT, RoBERTa and GPT-2 models, and its paper puts it at about 2.4 times the size of earlier Nepali corpora.

Openness

5 medium confidence
5.0
license
mit(Hugging Face card metadata only
access
public(Hugging Face gated: false)
dataset_card
present(brief card naming sources, article count and size)

The corpus downloads without a gate, and its card metadata tags it MIT, a software license, with no further terms for the scraped news and blog text. The card itself is a short summary, with collection details left to the paper.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single corpus repository. A download does not show whether the data was used to train a released model.

Capability

2 medium confidence
2.0

The corpus gives Nepali a crawled article collection larger than earlier Nepali corpora, described in a paper, and IRIIS Research trained its Nepali BERT, RoBERTa and GPT models on it, though no outside model is known to use it. At several billion tokens, estimated from its size on disk, it is far below the trillions of tokens in English pretraining corpora, and Sangraha, a documented corpus across twenty-two Indic languages, is far larger.

Verified 2026-09-24