IIT Bombay English-Hindi corpus
CFILT, IIT BombayThe IIT Bombay English-Hindi corpus is a parallel corpus of about 1.49 million sentence pairs for machine translation, with a companion monolingual Hindi corpus. It compiles existing public corpora such as OPUS subsets and HindEnCorp with material the Center for Indian Language Technology collected, including Indian government websites. It has served as the English-Hindi data for the Workshop on Asian Translation shared tasks since 2016.
Openness
2 high confidence- license
- cc-by-nc-4.0(LICENSE file and card badge
- access
- public(ungated on Hugging Face
- dataset_card
- present(card and project page list each source corpus with segment counts)
Anyone can download the corpus, but its license forbids commercial use, and the third-party corpora folded into it keep their original terms.
- https://huggingface.co/api/datasets/cfilt/iitb-english-hindi recorded 2026-09-24
gated: false; no license in cardData
- https://huggingface.co/datasets/cfilt/iitb-english-hindi/raw/main/README.md recorded 2026-09-24
License badge "CC BY-NC 4.0"; describes sources and WAT shared task use
- https://raw.githubusercontent.com/cfiltnlp/IITB-English-Hindi-PC/main/LICENSE recorded 2026-09-24
"Creative Commons Attribution-NonCommercial 4.0 International"
- https://www.cfilt.iitb.ac.in/iitb_parallel/ recorded 2026-09-24
"This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. The corpora we compiled from other sources are available under their respective licenses."
Adoption
2 high confidenceHugging Face downloads of the corpus repository. Downloads from the CFILT site and use through the WAT shared task are not counted.
- https://huggingface.co/api/datasets/cfilt/iitb-english-hindi recorded 2026-09-24
1646 downloads in the trailing 30 days for cfilt/iitb-english-hindi
Capability
2 medium confidenceThe IIT Bombay corpus was the largest public English-Hindi parallel corpus when published, is documented source by source, and has been the English-Hindi data for the Workshop on Asian Translation shared tasks. Its roughly one and a half million sentence pairs cover a single language pair, far from the billions of pairs in the largest parallel collections, and the Bharat Parallel Corpus Collection spans twenty-two Indian languages with far more pairs.
- https://arxiv.org/abs/1710.02855 recorded 2026-09-24
"The corpus contains 1.49 million parallel segments"; "used in two editions of shared tasks at the Workshop on Asian Language Translation"; "the largest publicly available English-Hindi parallel corpus"
- https://www.cfilt.iitb.ac.in/iitb_parallel/ recorded 2026-09-24
Source table: GNOME, KDE4, Tanzil, Tatoeba, OpenSubs2013 (Opus), HindEnCorp and others
Verified 2026-09-24