papers

Publications (32)

cs.CL2021

Semantic Search as Extractive Paraphrase Span Detection

Jenna Kanerva, Hanna Kitti, Li-Hsin Chang +3

In this paper, we approach the problem of semantic search by framing the search task as paraphrase span detection, i.e. given a segment of text as a query phrase, the task is to id…

q-bio.QM2016

An expanded evaluation of protein function prediction methods shows an improvement in accuracy

Yuxiang Jiang, Tal Ronnen Oron, Wyatt T Clark +144

Background: The increasing volume and variety of genotypic and phenotypic data is a major defining characteristic of modern biomedical sciences. At the same time, the limitations i…

cs.CL2019

Multilingual is not enough: BERT for Finnish

Antti Virtanen, Jenna Kanerva, Rami Ilo +5

Deep learning-based language models pretrained on large unannotated text corpora have been demonstrated to allow efficient transfer learning for natural language processing, with r…

cs.CL2022

Identifying gender bias in blockbuster movies through the lens of machine learning

Muhammad Junaid Haris, Aanchal Upreti, Melih Kurtaran +3

The problem of gender bias is highly prevalent and well known. In this paper, we have analysed the portrayal of gender roles in English movies, a medium that effectively influences…

cs.CL2023

Silver Syntax Pre-training for Cross-Domain Relation Extraction

Elisa Bassignana, Filip Ginter, Sampo Pyysalo +2

Relation Extraction (RE) remains a challenging task, especially when considering realistic out-of-domain evaluations. One of the main reasons for this is the limited training size…

cs.HC2025

Interaction Analysis by Humans and AI: A Comparative Perspective

Maryam Teimouri, Filip Ginter, Tomi "bgt" Suovuo

This paper explores how Mixed Reality (MR) and 2D video conferencing influence children's communication during a gesture-based guessing game. Finnish-speaking participants engaged…

cs.CL2021

Finnish Paraphrase Corpus

Jenna Kanerva, Filip Ginter, Li-Hsin Chang +7

In this paper, we introduce the first fully manually annotated paraphrase corpus for Finnish containing 53,572 paraphrase pairs harvested from alternative subtitles and news headin…

cs.CL2026

Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs

Joonatan Laato, Veera Schroderus, Jenna Kanerva +3

Digitized historical archives make it possible to study everyday social life on a large scale, but the information extracted directly from text often does not directly allow one to…

cs.CL2025

Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs

Joonatan Laato, Jenna Kanerva, John Loehr +2

We performed a zero-shot information extraction study on a historical collection of 89,339 brief Finnish-language interviews of refugee families relocated post-WWII from Finnish Ea…

cs.CL2023

FinGPT: Large Generative Models for a Small Language

Risto Luukkonen, Ville Komulainen, Jouni Luoma +18

Large language models (LLMs) excel in many tasks in NLP and beyond, but most open models have very limited coverage of smaller languages and LLM work tends to focus on languages wh…

cs.CL2025

OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches

Jenna Kanerva, Cassandra Ledins, Siiri Käpyaho +1

Optical Character Recognition (OCR) systems often introduce errors when transcribing historical documents, leaving room for post-correction to improve text quality. This study eval…

cs.CL2026

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19

Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…

cs.CL2021

Explaining Classes through Word Attribution

Samuel Rönnqvist, Amanda Myntti, Aki-Juhani Kyröläinen +3

In recent years, several methods have been proposed for explaining individual predictions of deep learning models, yet there has been little study of how to aggregate these predict…

cs.CL2021

Annotation Guidelines for the Turku Paraphrase Corpus

Jenna Kanerva, Filip Ginter, Li-Hsin Chang +8

This document describes the annotation guidelines used to construct the Turku Paraphrase Corpus. These guidelines were developed together with the corpus annotation, revising and e…

cs.CL2021

Quantitative Evaluation of Alternative Translations in a Corpus of Highly Dissimilar Finnish Paraphrases

Li-Hsin Chang, Sampo Pyysalo, Jenna Kanerva +1

In this paper, we present a quantitative evaluation of differences between alternative translations in a large recently released Finnish paraphrase corpus focusing in particular on…

cs.CL2019

Is Multilingual BERT Fluent in Language Generation?

Samuel Rönnqvist, Jenna Kanerva, Tapio Salakoski +1

The multilingual BERT model is trained on 104 languages and meant to serve as a universal language model and tool for encoding sentences. We explore how well the model performs on…

cs.CV2025

Creating a Historical Migration Dataset from Finnish Church Records, 1800-1920

Ari Vesalainen, Jenna Kanerva, Aida Nitsch +4

This article presents a large-scale effort to create a structured dataset of internal migration in Finland between 1800 and 1920 using digitized church moving records. These record…

cs.CL2020

Towards Fully Bilingual Deep Language Modeling

Li-Hsin Chang, Sampo Pyysalo, Jenna Kanerva +1

Language models based on deep neural networks have facilitated great advances in natural language processing and understanding tasks in recent years. While models covering a large…

cs.CL2019

Template-free Data-to-Text Generation of Finnish Sports News

Jenna Kanerva, Samuel Rönnqvist, Riina Kekki +2

News articles such as sports game reports are often thought to closely follow the underlying game statistics, but in practice they contain a notable amount of background knowledge,…

cs.CL2019

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

Nelda Kote, Marenglen Biba, Jenna Kanerva +2

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemm…

cs.CL2020

Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection

Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter +6

Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework.…

cs.CL2020

WikiBERT models: deep transfer learning for many languages

Sampo Pyysalo, Jenna Kanerva, Antti Virtanen +1

Deep neural language models such as BERT have enabled substantial recent advances in many natural language processing tasks. Due to the effort and computational cost involved in th…

cs.CL2020

Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks

Jenna Kanerva, Filip Ginter, Tapio Salakoski

In this paper we present a novel lemmatization method based on a sequence-to-sequence neural network architecture and morphosyntactic context representation. In the proposed method…

cs.CL2026

Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke

Yu Wu, Ananth Mahadevan, Filip Ginter +2

While digitized corpora have transformed the study of intellectual transmission, current methods rely heavily on lexical text reuse detection, capturing verbatim quotations but fun…

cs.CL2022

GEMv2: Multilingual NLG Benchmarking in a Single Line of Code

Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran +74

Evaluation in machine learning is usually informed by past choices, for example which datasets or metrics to use. This standardization enables the comparison on equal footing using…

cs.CL2025

FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering

Erik Henriksson, Otto Tarkka, Filip Ginter

Data quality is crucial for training Large Language Models (LLMs). Traditional heuristic filters often miss low-quality text or mistakenly remove valuable content. In this paper, w…

cs.CL2025

Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations

Emil Nuutinen, Iiro Rastas, Filip Ginter

We apply a simple method to machine translate datasets with span-level annotation using the DeepL MT service and its ability to translate formatted documents. Using this method, we…

cs.CL2019

Leveraging Text Repetitions and Denoising Autoencoders in OCR Post-correction

Kai Hakala, Aleksi Vesanto, Niko Miekka +2

A common approach for improving OCR quality is a post-processing step based on models correcting misdetected characters and tokens. These models are typically trained on aligned pa…

cs.CL2026

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

Amanda Myntti, Jenna Kanerva, Veronika Laippala +1

In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding models on five MTEB tasks sp…

cs.CL2021

Deep learning for sentence clustering in essay grading support

Li-Hsin Chang, Iiro Rastas, Sampo Pyysalo +1

Essays as a form of assessment test student knowledge on a deeper level than short answer and multiple-choice questions. However, the manual evaluation of essays is time- and labor…

cs.CL2023

Multi-CrossRE A Multi-Lingual Multi-Domain Dataset for Relation Extraction

Elisa Bassignana, Filip Ginter, Sampo Pyysalo +2

Most research in Relation Extraction (RE) involves the English language, mainly due to the lack of multi-lingual resources. We propose Multi-CrossRE, the broadest multi-lingual dat…

cs.CL2022

Out-of-Domain Evaluation of Finnish Dependency Parsing

Jenna Kanerva, Filip Ginter

The prevailing practice in the academia is to evaluate the model performance on in-domain evaluation data typically set aside from the training corpus. However, in many real world…