AI Potluck
Back to Gap Map Model components / Language-specific datasets

L3Cube-MahaCorpus

L3Cube Pune
restricted / Overall score: 1.0

L3Cube-MahaCorpus is a Marathi monolingual text corpus of 24.8 million sentences and 289 million tokens scraped from news and other websites. Combined with existing sources it forms a 752-million-token Marathi corpus, on which L3Cube Pune trained its MahaBERT, MahaAlBERT and MahaRoBERTa models and MahaFT word embeddings. Both versions are Google Drive downloads linked from the MarathiNLP repository.

Openness

2 medium confidence
2.0
license
cc-by-nc-sa-4.0(MarathiNLP README "License" section
access
public(direct Google Drive links in the README, no sign-up)
dataset_card
present(README section with news, non-news and full-corpus statistics

The README licenses MahaCorpus for non-commercial use with share-alike terms and describes it as released for research purposes. The files are plain Google Drive links with no sign-up.

Adoption

1 low confidence
1.0

GitHub stars on the MarathiNLP repository, which links the corpus alongside the group's Marathi models and other datasets, so the stars are not the corpus's alone; a star is attention rather than use.

Capability

1 medium confidence
1.0

MahaCorpus gives Marathi news and web sentences documented in a workshop paper, and the MahaBERT family of Marathi models was trained on it together with existing sources. At under three hundred million tokens it is far smaller than single-language sets such as the Basque Latxa Corpus and a tiny fraction of the trillions in English pretraining corpora.

  • https://arxiv.org/abs/2202.01159 recorded 2026-09-24

    Abstract: "We expand the existing Marathi monolingual corpus with 24.8M sentences and 289M tokens. We further present, MahaBERT, MahaAlBERT, and MahaRoBerta".

Verified 2026-09-24