OPUS logoOPUS logo
Loading search…
Contribute
OPUS APICorpus and language queriesMT APIEvaluation scores and modelsSynthetic APISynthetic collections and pairs
Data formatsDownload format referencePublicationsPapers and citations
CorporaBrowse released corporaSyntheticSynthetic corpus collectionsDashboardMT model scores and comparisons
Contribute
CorporaSyntheticDashboard
OPUS APIMT APISynthetic API
Data formatsPublications

Tools

Opus ToolsOpus FilterOPUS Distillery

Search

Opus QueryOpus WordalignOpus Explorer

Translation

OpusTranslate MobileAppOpusTranslate DesktopAppOPUS-CAT
Contribute to OPUSOpus LegacyGitHub
synOPUS logosynOPUS logo

Synthetic datasets

The synthetic open parallel corpus

Automatically generated and processed corpora for MT and LLM workflows.

synOPUS

synOPUS is a new edition that provides synthetic data sets — data that has (partially) been generated, for example, by translating text into other languages using machine translation tools or large language models. We used several tools to compile the current collection. All pre-processing is done automatically. No manual corrections have been carried out.

Released datasets

10
  • Europarlv8syn
  • Wikiv1syn
  • Wikibooksv1syn
  • Wikinewsv1syn
  • Wikipediav1syn
  • Wikiquotev1syn
  • Wikisourcev1syn
  • nemotron-cc-10K-samplev1syn
  • nemotron-cc-translatedv1syn
  • transweb-eduv1syn

SynOPUS tools

  • synOPUS Explorer
  • OpusTools
  • OpusFilter
  • OPUS-MT
  • OPUS-CAT
  • The Tatoeba Translation Challenge

All responses are available via HTTP endpoints and can be used from curl, Python, JavaScript, or any HTTP client.