papers

Publications (9)

cs.CL2018

Open Subtitles Paraphrase Corpus for Six Languages

Mathias Creutz

This paper accompanies the release of Opusparcus, a new paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish. The corpus consists of…

cs.CL2021

Semantic Search as Extractive Paraphrase Span Detection

Jenna Kanerva, Hanna Kitti, Li-Hsin Chang +3

In this paper, we approach the problem of semantic search by framing the search task as paraphrase span detection, i.e. given a segment of text as a query phrase, the task is to id…

cs.CL2021

Grammatical Error Generation Based on Translated Fragments

Eetu Sjöblom, Mathias Creutz, Teemu Vahtola

We perform neural machine translation of sentence fragments in order to create large amounts of training data for English grammatical error correction. Our method aims at simulatin…

cs.CL2023

On Using Distribution-Based Compositionality Assessment to Evaluate Compositional Generalisation in Machine Translation

Anssi Moisio, Mathias Creutz, Mikko Kurimo

Compositional generalisation (CG), in NLP and in machine learning more generally, has been assessed mostly using artificial datasets. It is important to develop benchmarks to asses…

cs.CL2024

LLMs' morphological analyses of complex FST-generated Finnish words

Anssi Moisio, Mathias Creutz, Mikko Kurimo

Rule-based language processing systems have been overshadowed by neural systems in terms of utility, but it remains unclear whether neural NLP systems, in practice, learn the gramm…

cs.CL2018

Paraphrase Detection on Noisy Subtitles in Six Languages

Eetu Sjöblom, Mathias Creutz, Mikko Aulamo

We perform automatic paraphrase detection on subtitle data from the Opusparcus corpus comprising six European languages: German, English, Finnish, French, Russian, and Swedish. We…

cs.CL2022

GEMv2: Multilingual NLG Benchmarking in a Single Line of Code

Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran +74

Evaluation in machine learning is usually informed by past choices, for example which datasets or metrics to use. This standardization enables the comparison on equal footing using…

cs.CL2002

Unsupervised Discovery of Morphemes

Mathias Creutz, Krista Lagus

We present two methods for unsupervised segmentation of words into morpheme-like units. The model utilized is especially suited for languages with a rich morphology, such as Finnis…

cs.CL2018

Multilingual NMT with a language-independent attention bridge

Raúl Vázquez, Alessandro Raganato, Jörg Tiedemann +1

In this paper, we propose a multilingual encoder-decoder architecture capable of obtaining multilingual sentence representations by means of incorporating an intermediate {\em atte…