AI Potluck
Back to Gap Map Organization

CFILT, IIT Bombay

lab · India

Scores

1 product on the map — 1 closed.

IIT Bombay English-Hindi corpus

Openness

2 high confidence
2.0
license
cc-by-nc-4.0(LICENSE file and card badge
access
public(ungated on Hugging Face
dataset_card
present(card and project page list each source corpus with segment counts)

Anyone can download the corpus, but its license forbids commercial use, and the third-party corpora folded into it keep their original terms.

Adoption

2 high confidence
2.0

Hugging Face downloads of the corpus repository. Downloads from the CFILT site and use through the WAT shared task are not counted.

Capability

2 medium confidence
2.0

The IIT Bombay corpus was the largest public English-Hindi parallel corpus when published, is documented source by source, and has been the English-Hindi data for the Workshop on Asian Translation shared tasks. Its roughly one and a half million sentence pairs cover a single language pair, far from the billions of pairs in the largest parallel collections, and the Bharat Parallel Corpus Collection spans twenty-two Indian languages with far more pairs.

  • https://arxiv.org/abs/1710.02855 recorded 2026-09-24

    "The corpus contains 1.49 million parallel segments"; "used in two editions of shared tasks at the Workshop on Asian Language Translation"; "the largest publicly available English-Hindi parallel corpus"

  • https://www.cfilt.iitb.ac.in/iitb_parallel/ recorded 2026-09-24

    Source table: GNOME, KDE4, Tanzil, Tatoeba, OpenSubs2013 (Opus), HindEnCorp and others