CFILT, IIT Bombay
lab · IndiaScores
1 product on the map — 1 closed.
IIT Bombay English-Hindi corpus
Openness
2 high confidence- license
- cc-by-nc-4.0(LICENSE file and card badge
- access
- public(ungated on Hugging Face
- dataset_card
- present(card and project page list each source corpus with segment counts)
Anyone can download the corpus, but its license forbids commercial use, and the third-party corpora folded into it keep their original terms.
- https://huggingface.co/api/datasets/cfilt/iitb-english-hindi recorded 2026-09-24
gated: false; no license in cardData
- https://huggingface.co/datasets/cfilt/iitb-english-hindi/raw/main/README.md recorded 2026-09-24
License badge "CC BY-NC 4.0"; describes sources and WAT shared task use
- https://raw.githubusercontent.com/cfiltnlp/IITB-English-Hindi-PC/main/LICENSE recorded 2026-09-24
"Creative Commons Attribution-NonCommercial 4.0 International"
- https://www.cfilt.iitb.ac.in/iitb_parallel/ recorded 2026-09-24
"This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. The corpora we compiled from other sources are available under their respective licenses."
Adoption
2 high confidenceHugging Face downloads of the corpus repository. Downloads from the CFILT site and use through the WAT shared task are not counted.
- https://huggingface.co/api/datasets/cfilt/iitb-english-hindi recorded 2026-09-24
1646 downloads in the trailing 30 days for cfilt/iitb-english-hindi
Capability
2 medium confidenceThe IIT Bombay corpus was the largest public English-Hindi parallel corpus when published, is documented source by source, and has been the English-Hindi data for the Workshop on Asian Translation shared tasks. Its roughly one and a half million sentence pairs cover a single language pair, far from the billions of pairs in the largest parallel collections, and the Bharat Parallel Corpus Collection spans twenty-two Indian languages with far more pairs.
- https://arxiv.org/abs/1710.02855 recorded 2026-09-24
"The corpus contains 1.49 million parallel segments"; "used in two editions of shared tasks at the Workshop on Asian Language Translation"; "the largest publicly available English-Hindi parallel corpus"
- https://www.cfilt.iitb.ac.in/iitb_parallel/ recorded 2026-09-24
Source table: GNOME, KDE4, Tanzil, Tatoeba, OpenSubs2013 (Opus), HindEnCorp and others