Pralekha
AI4BharatPralekha is a parallel corpus and benchmark of whole documents covering 11 Indic languages and English, with more than 3 million aligned document pairs, 1.5 million of them English-Indic. The documents are Press Information Bureau releases and Mann Ki Baat radio scripts, and unalignable documents sampled from Sangraha are mixed in to test cross-lingual document alignment. AI4Bharat built it, with train, dev and test splits for translating whole documents.
Openness
5 high confidence- license
- cc-by-4.0(card metadata and license badge)
- access
- public
- dataset_card
- present(card describes domains, subsets, splits and per-language counts)
The documents are released under an attribution-only license with no gate, and the alignment code is public in the GitHub repository.
- https://huggingface.co/api/datasets/ai4bharat/Pralekha recorded 2026-09-24
gated: false; cardData license: cc-by-4.0
- https://huggingface.co/datasets/ai4bharat/Pralekha/raw/main/README.md recorded 2026-09-24
license: cc-by-4.0; License badge "CC BY 4.0"; alignable and unalignable subsets with language-wise statistics
- https://raw.githubusercontent.com/AI4Bharat/Pralekha/master/README.md recorded 2026-09-24
README for the alignment pipeline: setup, input directory structure and usage
Adoption
1 high confidenceHugging Face downloads of the corpus repository. Clones of the alignment code on GitHub are not counted.
- https://huggingface.co/api/datasets/ai4bharat/Pralekha recorded 2026-09-24
522 downloads in the trailing 30 days for ai4bharat/Pralekha
Capability
2 medium confidencePralekha aligns documents from government releases and radio scripts across eleven Indic languages and English, documented in a paper and card, mainly to benchmark document alignment with translation training as a second use, though no model outside its paper names it. At a few million document pairs it is well short of the billions of pairs in the largest parallel collections, and the Bharat Parallel Corpus Collection underlies the IndicTrans translation models.
- https://arxiv.org/abs/2411.19096 recorded 2026-09-24
"Pralekha, a benchmark containing over 3 million aligned document pairs across 11 Indic languages and English, which includes 1.5 million English-Indic pairs"
Verified 2026-09-24