AI Potluck
Back to Gap Map Model components / Language-specific datasets

IndicXTREME

AI4Bharat
restricted / Overall score: 3.4

IndicXTREME is a natural language understanding benchmark of nine tasks across 20 Indian languages, grouping 105 evaluation sets for classification, structure prediction, question answering and sentence retrieval. AI4Bharat created new sets, among them IndicCOPA, IndicQA, IndicXParaphrase and IndicSentiment, and folded in Naamapadam, IndicXNLI, MASSIVE and FLORES. It was released with IndicBERT v2 to test zero-shot multilingual transfer.

Listed are the AI4Bharat-hosted task repositories; the IndicXNLI, MASSIVE and FLORES tasks belong to other publishers and are not counted here.

Openness

2 medium confidence
2.0
license
cc-by-4.0(IndicCOPA and IndicQA metadata)+cc0-1.0(Naamapadam metadata
the IndicBERT README says all datasets from the work are CC-0)+not-clearly-stated-on-card(IndicXParaphrase and IndicSentiment
no license in card or API metadata)
access
public(all five repositories ungated)
dataset_card
partial(IndicCOPA card is an empty template and IndicXParaphrase has none

The test sets download without a gate and the project README dedicates the new sets to the public domain. The Hugging Face repositories disagree with it, two carrying attribution licenses and two carrying none, and most per-task cards are empty.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the five AI4Bharat task repositories. Naamapadam downloads include use of its training set outside the benchmark, and the third-party tasks are not counted.

Capability

4 medium confidence
4.0

IndicXTREME gathers nine understanding tasks across 20 Indian languages, and IndicBERT v2 and the Airavata evaluation suite report against it. It is narrower than AfroBench, which spans 15 tasks and 64 languages behind a public leaderboard, and some of its sets are machine-translated.

Verified 2026-09-24