Institute for Research and Innovation in Intelligent Systems (IRIIS)
labScores
1 product on the map — 1 open.
Openness
5 medium confidence- license
- mit(Hugging Face card metadata only
- access
- public(Hugging Face gated: false)
- dataset_card
- present(brief card naming sources, article count and size)
The corpus downloads without a gate, and its card metadata tags it MIT, a software license, with no further terms for the scraped news and blog text. The card itself is a short summary, with collection details left to the paper.
- https://huggingface.co/api/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus recorded 2026-09-24
"gated": false; cardData license "mit".
- https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus recorded 2026-09-24
"Total Articles: ~6.4 million; Language: Nepali; Size: 27.5 GB (in csv); Source: Collected from various Nepali news websites, blogs, and other online platforms."
Adoption
1 high confidenceHugging Face downloads of the single corpus repository. A download does not show whether the data was used to train a released model.
- https://huggingface.co/api/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus recorded 2026-09-24
996 downloads in the trailing 30 days for IRIIS-RESEARCH/Nepali-Text-Corpus
Capability
2 medium confidenceThe corpus gives Nepali a crawled article collection larger than earlier Nepali corpora, described in a paper, and IRIIS Research trained its Nepali BERT, RoBERTa and GPT models on it, though no outside model is known to use it. At several billion tokens, estimated from its size on disk, it is far below the trillions of tokens in English pretraining corpora, and Sangraha, a documented corpus across twenty-two Indic languages, is far larger.
- https://arxiv.org/abs/2411.15734 recorded 2026-09-24
Abstract: "we have collected 27.5 GB of Nepali text data, approximately 2.4x larger than any previously available Nepali language corpus".
- https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus recorded 2026-09-24
"Models trained or fine-tuned on" lists IRIIS-RESEARCH/RoBERTa_Nepali_125M and dinesh-bk/NepGPT2.