AI Potluck
Back to Gap Map Model components / Language-specific datasets

Pralekha

AI4Bharat
open / Overall score: 1.7

Pralekha is a parallel corpus and benchmark of whole documents covering 11 Indic languages and English, with more than 3 million aligned document pairs, 1.5 million of them English-Indic. The documents are Press Information Bureau releases and Mann Ki Baat radio scripts, and unalignable documents sampled from Sangraha are mixed in to test cross-lingual document alignment. AI4Bharat built it, with train, dev and test splits for translating whole documents.

Openness

5 high confidence
5.0
license
cc-by-4.0(card metadata and license badge)
access
public
dataset_card
present(card describes domains, subsets, splits and per-language counts)

The documents are released under an attribution-only license with no gate, and the alignment code is public in the GitHub repository.

Adoption

1 high confidence
1.0

Hugging Face downloads of the corpus repository. Clones of the alignment code on GitHub are not counted.

Capability

2 medium confidence
2.0

Pralekha aligns documents from government releases and radio scripts across eleven Indic languages and English, documented in a paper and card, mainly to benchmark document alignment with translation training as a second use, though no model outside its paper names it. At a few million document pairs it is well short of the billions of pairs in the largest parallel collections, and the Bharat Parallel Corpus Collection underlies the IndicTrans translation models.

  • https://arxiv.org/abs/2411.19096 recorded 2026-09-24

    "Pralekha, a benchmark containing over 3 million aligned document pairs across 11 Indic languages and English, which includes 1.5 million English-Indic pairs"

Verified 2026-09-24