Publications (9)
Open Subtitles Paraphrase Corpus for Six Languages
Mathias Creutz
This paper accompanies the release of Opusparcus, a new paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish. The corpus consists of…
Semantic Search as Extractive Paraphrase Span Detection
Jenna Kanerva, Hanna Kitti, Li-Hsin Chang +3
In this paper, we approach the problem of semantic search by framing the search task as paraphrase span detection, i.e. given a segment of text as a query phrase, the task is to id…
Grammatical Error Generation Based on Translated Fragments
Eetu Sjöblom, Mathias Creutz, Teemu Vahtola
We perform neural machine translation of sentence fragments in order to create large amounts of training data for English grammatical error correction. Our method aims at simulatin…
On Using Distribution-Based Compositionality Assessment to Evaluate Compositional Generalisation in Machine Translation
Anssi Moisio, Mathias Creutz, Mikko Kurimo
Compositional generalisation (CG), in NLP and in machine learning more generally, has been assessed mostly using artificial datasets. It is important to develop benchmarks to asses…
LLMs' morphological analyses of complex FST-generated Finnish words
Anssi Moisio, Mathias Creutz, Mikko Kurimo
Rule-based language processing systems have been overshadowed by neural systems in terms of utility, but it remains unclear whether neural NLP systems, in practice, learn the gramm…
Paraphrase Detection on Noisy Subtitles in Six Languages
Eetu Sjöblom, Mathias Creutz, Mikko Aulamo
We perform automatic paraphrase detection on subtitle data from the Opusparcus corpus comprising six European languages: German, English, Finnish, French, Russian, and Swedish. We…
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran +74
Evaluation in machine learning is usually informed by past choices, for example which datasets or metrics to use. This standardization enables the comparison on equal footing using…
Unsupervised Discovery of Morphemes
Mathias Creutz, Krista Lagus
We present two methods for unsupervised segmentation of words into morpheme-like units. The model utilized is especially suited for languages with a rich morphology, such as Finnis…
Multilingual NMT with a language-independent attention bridge
Raúl Vázquez, Alessandro Raganato, Jörg Tiedemann +1
In this paper, we propose a multilingual encoder-decoder architecture capable of obtaining multilingual sentence representations by means of incorporating an intermediate {\em atte…