OPUS logoOPUS logo
Loading search…
Contribute
OPUS APICorpus and language queriesMT APIEvaluation scores and modelsSynthetic APISynthetic collections and pairs
Data formatsDownload format referencePublicationsPapers and citations
CorporaBrowse released corporaSyntheticSynthetic corpus collectionsDashboardMT model scores and comparisons
Contribute
CorporaSyntheticDashboard
OPUS APIMT APISynthetic API
Data formatsPublications

Tools

Opus ToolsOpus FilterOPUS Distillery

Search

Opus QueryOpus WordalignOpus Explorer

Translation

OpusTranslate MobileAppOpusTranslate DesktopAppOPUS-CAT
Contribute to OPUSOpus LegacyGitHub
Download

IITBv2.0

Latest releasev2.0
LicenseCC BY-NC 4.0

Copyright

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. The corpora we compiled from other sources are available under their respective licenses. More information is available in https://www.cfilt.iitb.ac.in/~moses/iitb_en_hi_parallel/
The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task in 2016 and 2017 for the Hindi-to-English and English-to-Hindi languages pairs and as a pivot language pair for the Hindi-to-Japanese and Japanese-to-Hindi language pairs.
If you use this corpus or its derivate resources for your research, kindly cite it as follows: Anoop Kunchukuttan, Pratik Mehta, Pushpak Bhattacharyya. The IIT Bombay English-Hindi Parallel Corpus. Language Resources and Evaluation Conference. 2018.

IITB's Numbers

LanguagesBitextsFilesTokensSentence fragments
21256.80M3.31M

Downloads

Available resources

This corpus has one language pair, so the resources are listed directly.

Available language pair

English (en) - Hindi (hi)

v2.0

Sample
Sentences
1,561,173
en tokens
22,594,637
hi tokens
34,202,009

Bilingual downloads

Monolingual downloads

en
hi

TMX files contain unique translation units. Moses downloads include all non-empty alignment units, including duplicates.

Disclaimer

  • We do not own any of the text from which the data has been extracted.
  • We only offer files that we believe we are free to redistribute. If any doubt occurs about the legality of any of our file downloads we will take them off right away after contacting us.

Notice and take down policy

Notice: Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please:

  • Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted.
  • Clearly identify the copyrighted work claimed to be infringed.
  • Clearly identify the material that is claimed to be infringing and information reasonably sufficient to allow us to locate the material.
  • And contact the OPUS project at: opus-project at helsinki.fi.

Take down: We will comply to legitimate requests by removing the affected sources from the next release of the corpus.