AI Potluck
Back to Gap Map Model components / Language-specific datasets

Bharat Parallel Corpus Collection

AI4Bharat
gated / Overall score: 3.1

The Bharat Parallel Corpus Collection (BPCC) is an English-Indic parallel text corpus of about 230 million sentence pairs covering all 22 scheduled Indian languages. Most pairs are mined from the web or drawn from corpora such as Samanantar and NLLB, while BPCC-Human holds about 2.2 million human-translated pairs, including new Wikipedia and everyday-use sets. AI4Bharat released it to train IndicTrans2, together with back-translated data.

The IN22 test sets released with BPCC live in their own repositories and are not scored here. The IndicTrans2 README says newer BPCC-Seed releases are published on the AI4Bharat website.

Openness

3 high confidence
3.0
license
cc0-1.0(mined corpora, existing seed corpora (NLLB-Seed, ILCI, MASSIVE) and back-translation data)+cc-by-4.0(new human-translated BPCC-H-Wiki and BPCC-H-Daily sets)
access
auto(Hugging Face gated: auto)
dataset_card
present(card lists every source and subset with pair counts and licenses)

AI4Bharat waives its rights to the mined, existing and back-translated parts and licenses its new human translations with attribution, while stating that it does not own the underlying source text. The Hugging Face copy needs an access request that is approved automatically.

Adoption

1 high confidence
1.0

Hugging Face downloads of the BPCC repository. The IndicTrans2 README also links a full zip of the data from the same repository, and newer seed data from the AI4Bharat website is not counted.

Capability

4 medium confidence
4.0

The Bharat Parallel Corpus Collection is the largest public parallel corpus for Indian languages, mostly mined from the web around a human-translated core of a couple of million pairs, and its builders trained IndicTrans, the first translation model to cover all twenty-two scheduled languages, on it. Its few hundred million sentence pairs are still a fraction of the largest parallel collections, which run to billions of pairs.

  • https://arxiv.org/abs/2305.16307 recorded 2026-09-24

    Abstract: BPCC is "the largest publicly available parallel corpora for Indic languages. BPCC contains a total of 230M bitext pairs"; IndicTrans2 is "the first model to support all 22 languages".

Verified 2026-09-24