Aksharantar
AI4BharatAksharantar is a transliteration corpus of about 26 million word pairs between English romanization and 21 Indic languages written in 12 scripts. Most pairs were mined from monolingual and parallel corpora such as IndicCorp and Samanantar, and human annotators added transliterated words and named entities, plus a test set of 103,000 pairs. AI4Bharat built it alongside the IndicXlit transliteration model.
Openness
5 high confidence- license
- cc-by(manually collected data
- access
- public
- dataset_card
- present(card gives languages, fields, splits and licensing
The mined pairs are dedicated to the public domain and the hand-collected pairs need only attribution, so every part can be reused commercially. The files download without any gate.
- https://huggingface.co/api/datasets/ai4bharat/Aksharantar recorded 2026-09-24
gated: false; cardData license: cc
- https://huggingface.co/datasets/ai4bharat/Aksharantar/raw/main/README.md recorded 2026-09-24
"Manually collected data: Released under CC-BY license. Mined dataset (from Samanantar and IndicCorp): Released under CC0 license. Existing sources: Released under CC0 license."
Adoption
1 high confidenceHugging Face downloads of the corpus repository. Most use is likely through the IndicXlit model trained on it, which these downloads do not count.
- https://huggingface.co/api/datasets/ai4bharat/Aksharantar recorded 2026-09-24
685 downloads in the trailing 30 days for ai4bharat/Aksharantar
Capability
3 medium confidenceAksharantar is the largest public transliteration resource for Indian languages, many times earlier sets, and the IndicXlit model is trained and evaluated on it. Its tens of millions of word pairs serve one task, well short of the billions of pairs in the largest English and Chinese parallel collections, and the Bharat Parallel Corpus Collection from the same team covers full translation across all twenty-two scheduled languages.
- https://arxiv.org/abs/2205.03018 recorded 2026-09-24
"the largest publicly available transliteration dataset for Indian languages"; "26 million transliteration pairs for 21 Indic languages"; "21 times larger than existing datasets"; "we trained IndicXlit"
- https://raw.githubusercontent.com/AI4Bharat/IndicXlit/master/README.md recorded 2026-09-24
IndicXlit "is trained on Aksharantar" and "is evaluated on Dakshina benchmark and Aksharantar benchmark"
Verified 2026-09-24