OPUS reference
OPUS publishes corpus data in a few related formats. The native representation is XML; plain text, Moses, TMX, and frequency downloads are derived from those corpus files.
Native corpus data
OPUS corpus files are organized by corpus, processing level, and language. Untokenized files live in the raw layer, while tokenized XML files live in the XML layer. Sentence boundaries are kept compatible across those layers so they can be connected by alignment files.
Use this when you want basic XML markup and original text as closely as possible after corpus segmentation.
Use this when you want XML files where tokens are represented explicitly, usually with token-level markup.
Corpus/raw/en/example.xml.gz Corpus/xml/en/example.xml.gz Corpus/xml/en-fr.xml.gz
Sentence alignment
Bilingual XML downloads contain standoff sentence alignments in XCES Align format. They reference source and target documents and connect sentence IDs with alignment links. Only one direction is stored for a language pair, so the filename order may be normalized rather than matching the direction you selected.
<linkGrp fromDoc="en/doc.xml.gz" toDoc="fr/doc.xml.gz"> <link xtargets="1;1" /> <link xtargets="2;2 3" /> </linkGrp>
Plain text
Plain text downloads are convenient when you want line-based data. Monolingual text files contain one sentence per line. Moses bitext downloads contain two files whose lines are aligned with each other; empty alignments are excluded.
Corpus.en-fr.en Corpus.en-fr.fr
Untokenized plain text for a single language.
Tokenized plain text for a single language.
Translation memory
TMX files package bilingual data as translation memory units. In OPUS they use minimal markup and raw text segments, making them useful for translation-memory tools and exchange workflows.
<tu> <tuv xml:lang="en"><seg>One source segment.</seg></tuv> <tuv xml:lang="fr"><seg>One target segment.</seg></tuv> </tu>