IndicXTREME
AI4BharatIndicXTREME is a natural language understanding benchmark of nine tasks across 20 Indian languages, grouping 105 evaluation sets for classification, structure prediction, question answering and sentence retrieval. AI4Bharat created new sets, among them IndicCOPA, IndicQA, IndicXParaphrase and IndicSentiment, and folded in Naamapadam, IndicXNLI, MASSIVE and FLORES. It was released with IndicBERT v2 to test zero-shot multilingual transfer.
Listed are the AI4Bharat-hosted task repositories; the IndicXNLI, MASSIVE and FLORES tasks belong to other publishers and are not counted here.
Openness
2 medium confidence- license
- cc-by-4.0(IndicCOPA and IndicQA metadata)+cc0-1.0(Naamapadam metadata
- the IndicBERT README says all datasets from the work are CC-0)+not-clearly-stated-on-card(IndicXParaphrase and IndicSentiment
- no license in card or API metadata)
- access
- public(all five repositories ungated)
- dataset_card
- partial(IndicCOPA card is an empty template and IndicXParaphrase has none
The test sets download without a gate and the project README dedicates the new sets to the public domain. The Hugging Face repositories disagree with it, two carrying attribution licenses and two carrying none, and most per-task cards are empty.
- https://huggingface.co/api/datasets/ai4bharat/IndicCOPA recorded 2026-09-24
gated: false; cardData license: [cc-by-4.0]
- https://huggingface.co/api/datasets/ai4bharat/IndicQA recorded 2026-09-24
gated: false; cardData license: [cc-by-4.0]
- https://huggingface.co/api/datasets/ai4bharat/IndicSentiment recorded 2026-09-24
gated: false; no license in cardData
- https://huggingface.co/api/datasets/ai4bharat/IndicXParaphrase recorded 2026-09-24
gated: false; no license in cardData
- https://huggingface.co/api/datasets/ai4bharat/naamapadam recorded 2026-09-24
gated: false; cardData license: [cc0-1.0]
- https://huggingface.co/datasets/ai4bharat/IndicCOPA/raw/main/README.md recorded 2026-09-24
Card template with "[More Information Needed]" in every section
- https://raw.githubusercontent.com/AI4Bharat/IndicBERT/main/README.md recorded 2026-09-24
"All the datasets created as part of this work will be released under a CC-0 license"; task list with dataset links
Adoption
2 high confidenceHugging Face downloads summed over the five AI4Bharat task repositories. Naamapadam downloads include use of its training set outside the benchmark, and the third-party tasks are not counted.
- https://huggingface.co/api/datasets/ai4bharat/IndicCOPA recorded 2026-09-24
425 downloads in the trailing 30 days for ai4bharat/IndicCOPA
- https://huggingface.co/api/datasets/ai4bharat/IndicQA recorded 2026-09-24
1708 downloads in the trailing 30 days for ai4bharat/IndicQA
- https://huggingface.co/api/datasets/ai4bharat/IndicSentiment recorded 2026-09-24
612 downloads in the trailing 30 days for ai4bharat/IndicSentiment
- https://huggingface.co/api/datasets/ai4bharat/IndicXParaphrase recorded 2026-09-24
362 downloads in the trailing 30 days for ai4bharat/IndicXParaphrase
- https://huggingface.co/api/datasets/ai4bharat/naamapadam recorded 2026-09-24
770 downloads in the trailing 30 days for ai4bharat/naamapadam
Capability
4 medium confidenceIndicXTREME gathers nine understanding tasks across 20 Indian languages, and IndicBERT v2 and the Airavata evaluation suite report against it. It is narrower than AfroBench, which spans 15 tasks and 64 languages behind a public leaderboard, and some of its sets are machine-translated.
- https://arxiv.org/abs/2212.10168 recorded 2026-09-24
Naamapadam: training data "automatically created from the Samanantar parallel corpus"; "manually annotated testsets for 9 languages"
- https://huggingface.co/datasets/ai4bharat/IndicCOPA recorded 2026-09-24
Collections including ai4bharat/IndicCOPA: "Airavata Evaluation Suite" and "IndicXTREME"
- https://raw.githubusercontent.com/AI4Bharat/IndicBERT/main/README.md recorded 2026-09-24
"IndicXTREME benchmark includes 9 tasks"; "a human-supervised benchmark, IndicXTREME, consisting of nine diverse NLU tasks covering 20 languages ... 105 evaluation sets, of which 52 are new"
Verified 2026-09-24