OPUS logoOPUS logo
Loading search…
Contribute
OPUS APICorpus and language queriesMT APIEvaluation scores and modelsSynthetic APISynthetic collections and pairs
Data formatsDownload format referencePublicationsPapers and citations
CorporaBrowse released corporaSyntheticSynthetic corpus collectionsDashboardMT model scores and comparisons
Contribute
CorporaSyntheticDashboard
OPUS APIMT APISynthetic API
Data formatsPublications

Tools

Opus ToolsOpus FilterOPUS Distillery

Search

Opus QueryOpus WordalignOpus Explorer

Translation

OpusTranslate MobileAppOpusTranslate DesktopAppOPUS-CAT
Contribute to OPUSOpus LegacyGitHub

OPUS reference

Download formats

OPUS publishes corpus data in a few related formats. The native representation is XML; plain text, Moses, TMX, and frequency downloads are derived from those corpus files.

Browse corporaLegacy reference

Monolingual downloads

xml-raw
Basic XML-encoded corpus files.
xml-tok
Tokenized XML-encoded corpus files.
txt-raw
Raw plain text files, one sentence per line.
txt-tok
Tokenized plain text files, one sentence per line.
freq
Token frequency lists when available.

Bilingual downloads

moses
Aligned plain text files.
TMX
Translation memories in a standard exchange format.
XML
Sentence alignments in XCES Align format.

Native corpus data

XML files

OPUS corpus files are organized by corpus, processing level, and language. Untokenized files live in the raw layer, while tokenized XML files live in the XML layer. Sentence boundaries are kept compatible across those layers so they can be connected by alignment files.

xml-raw

Use this when you want basic XML markup and original text as closely as possible after corpus segmentation.

xml-tok

Use this when you want XML files where tokens are represented explicitly, usually with token-level markup.

Corpus/raw/en/example.xml.gz
Corpus/xml/en/example.xml.gz
Corpus/xml/en-fr.xml.gz

Sentence alignment

Bilingual XML

Bilingual XML downloads contain standoff sentence alignments in XCES Align format. They reference source and target documents and connect sentence IDs with alignment links. Only one direction is stored for a language pair, so the filename order may be normalized rather than matching the direction you selected.

<linkGrp fromDoc="en/doc.xml.gz" toDoc="fr/doc.xml.gz">
  <link xtargets="1;1" />
  <link xtargets="2;2 3" />
</linkGrp>

Plain text

Text and Moses files

Plain text downloads are convenient when you want line-based data. Monolingual text files contain one sentence per line. Moses bitext downloads contain two files whose lines are aligned with each other; empty alignments are excluded.

Corpus.en-fr.en
Corpus.en-fr.fr

txt-raw

Untokenized plain text for a single language.

txt-tok

Tokenized plain text for a single language.

Translation memory

TMX

TMX files package bilingual data as translation memory units. In OPUS they use minimal markup and raw text segments, making them useful for translation-memory tools and exchange workflows.

<tu>
  <tuv xml:lang="en"><seg>One source segment.</seg></tuv>
  <tuv xml:lang="fr"><seg>One target segment.</seg></tuv>
</tu>