L3Cube-MahaCorpus
L3Cube PuneL3Cube-MahaCorpus is a Marathi monolingual text corpus of 24.8 million sentences and 289 million tokens scraped from news and other websites. Combined with existing sources it forms a 752-million-token Marathi corpus, on which L3Cube Pune trained its MahaBERT, MahaAlBERT and MahaRoBERTa models and MahaFT word embeddings. Both versions are Google Drive downloads linked from the MarathiNLP repository.
Openness
2 medium confidence- license
- cc-by-nc-sa-4.0(MarathiNLP README "License" section
- access
- public(direct Google Drive links in the README, no sign-up)
- dataset_card
- present(README section with news, non-news and full-corpus statistics
The README licenses MahaCorpus for non-commercial use with share-alike terms and describes it as released for research purposes. The files are plain Google Drive links with no sign-up.
- https://raw.githubusercontent.com/l3cube-pune/MarathiNLP/main/README.md recorded 2026-09-24
"L3Cube-MahaCorpus ... are licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. The datasets are released to the community for research purposes only".
- https://raw.githubusercontent.com/l3cube-pune/MarathiNLP/main/README.md recorded 2026-09-24
Table: MahaCorpus (news) 212M tokens, (non-news) 76.4M, (full) 289M / 24.8M sentences, Full Marathi Corpus 752M, each with a drive.google.com link.
Adoption
1 low confidenceGitHub stars on the MarathiNLP repository, which links the corpus alongside the group's Marathi models and other datasets, so the stars are not the corpus's alone; a star is attention rather than use.
- https://ungh.cc/repos/l3cube-pune/MarathiNLP recorded 2026-09-24
GitHub repository record for l3cube-pune/MarathiNLP (via the ungh.cc mirror of the GitHub API): 164 stargazers, last push 2025-09-14.
Capability
1 medium confidenceMahaCorpus gives Marathi news and web sentences documented in a workshop paper, and the MahaBERT family of Marathi models was trained on it together with existing sources. At under three hundred million tokens it is far smaller than single-language sets such as the Basque Latxa Corpus and a tiny fraction of the trillions in English pretraining corpora.
- https://arxiv.org/abs/2202.01159 recorded 2026-09-24
Abstract: "We expand the existing Marathi monolingual corpus with 24.8M sentences and 289M tokens. We further present, MahaBERT, MahaAlBERT, and MahaRoBerta".
Verified 2026-09-24