AI Potluck
Back to Gap Map Model components / Language-specific datasets

IndicCorp v2

AI4Bharat
open / Overall score: 2.7

IndicCorp v2 is a monolingual pretraining text corpus of about 20.9 billion tokens in 24 languages from four language families, reaching low-resource languages such as Bodo, Khasi and Santali. AI4Bharat curated it to train its IndicBERT v2 encoder models, and released it with IndicXTREME, a human-supervised benchmark of nine language-understanding tasks.

Openness

4 medium confidence
4.0
license
cc0-1.0(card and IndicBERT README: "All the datasets created as part of this work will be released under a CC-0 license"
access
public(Hugging Face gated: false)
dataset_card
partial(card gives usage, license and citation but no source or size breakdown

The card and README release the corpus under CC0 and it downloads without a gate. The card itself only shows how to load it, and what the corpus contains is described only in the paper.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single corpus repository, fetched per language split. A download does not show whether the data went into a released model.

Capability

3 medium confidence
3.0

AI4Bharat trained IndicBERT v2 on IndicCorp v2, and its newer IndicBERT v3 and Cadence models still list it as training data. The same team's newer Sangraha is about twelve times larger, which makes IndicCorp v2 the superseded one of AI4Bharat's two pretraining corpora.

Verified 2026-09-24