AI Potluck
Back to Gap Map Model components / Language-specific datasets

Aksharantar

AI4Bharat
open / Overall score: 2.4

Aksharantar is a transliteration corpus of about 26 million word pairs between English romanization and 21 Indic languages written in 12 scripts. Most pairs were mined from monolingual and parallel corpora such as IndicCorp and Samanantar, and human annotators added transliterated words and named entities, plus a test set of 103,000 pairs. AI4Bharat built it alongside the IndicXlit transliteration model.

Openness

5 high confidence
5.0
license
cc-by(manually collected data
access
public
dataset_card
present(card gives languages, fields, splits and licensing

The mined pairs are dedicated to the public domain and the hand-collected pairs need only attribution, so every part can be reused commercially. The files download without any gate.

Adoption

1 high confidence
1.0

Hugging Face downloads of the corpus repository. Most use is likely through the IndicXlit model trained on it, which these downloads do not count.

Capability

3 medium confidence
3.0

Aksharantar is the largest public transliteration resource for Indian languages, many times earlier sets, and the IndicXlit model is trained and evaluated on it. Its tens of millions of word pairs serve one task, well short of the billions of pairs in the largest English and Chinese parallel collections, and the Bharat Parallel Corpus Collection from the same team covers full translation across all twenty-two scheduled languages.

Verified 2026-09-24