IndicCorp v2
AI4BharatIndicCorp v2 is a monolingual pretraining text corpus of about 20.9 billion tokens in 24 languages from four language families, reaching low-resource languages such as Bodo, Khasi and Santali. AI4Bharat curated it to train its IndicBERT v2 encoder models, and released it with IndicXTREME, a human-supervised benchmark of nine language-understanding tasks.
Openness
4 medium confidence- license
- cc0-1.0(card and IndicBERT README: "All the datasets created as part of this work will be released under a CC-0 license"
- access
- public(Hugging Face gated: false)
- dataset_card
- partial(card gives usage, license and citation but no source or size breakdown
The card and README release the corpus under CC0 and it downloads without a gate. The card itself only shows how to load it, and what the corpus contains is described only in the paper.
- https://huggingface.co/api/datasets/ai4bharat/IndicCorpV2 recorded 2026-09-24
"gated": false; no license tag.
- https://huggingface.co/datasets/ai4bharat/IndicCorpV2 recorded 2026-09-24
"All the datasets created as part of this work will be released under a CC-0 license and all models & code will be release under an MIT license"; card shows loading example and citation.
- https://raw.githubusercontent.com/AI4Bharat/IndicBERT/main/README.md recorded 2026-09-24
LICENSE section repeats the CC-0 statement; "The dataset is now available on HuggingFace".
Adoption
2 high confidenceHugging Face downloads of the single corpus repository, fetched per language split. A download does not show whether the data went into a released model.
- https://huggingface.co/api/datasets/ai4bharat/IndicCorpV2 recorded 2026-09-24
3945 downloads in the trailing 30 days for ai4bharat/IndicCorpV2
Capability
3 medium confidenceAI4Bharat trained IndicBERT v2 on IndicCorp v2, and its newer IndicBERT v3 and Cadence models still list it as training data. The same team's newer Sangraha is about twelve times larger, which makes IndicCorp v2 the superseded one of AI4Bharat's two pretraining corpora.
- https://arxiv.org/abs/2212.05409 recorded 2026-09-24
Abstract: "we curate the largest monolingual corpora, IndicCorp, with 20.9B tokens covering 24 languages from 4 language families".
- https://huggingface.co/datasets/ai4bharat/IndicCorpV2 recorded 2026-09-24
"Models trained or fine-tuned on ai4bharat/IndicCorpV2" lists Cadence-Fast, Cadence, IndicBERT-v3-270M, IndicBERT-v3-4B, IndicBERT-v3-1B.
- https://huggingface.co/datasets/ai4bharat/Kathbath recorded 2026-09-24
"The text data comes from the IndicCorp dataset which is a crawl of publicly available websites."
Verified 2026-09-24