collaborators

6 papers

cs.CL2026

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19

Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…

cs.CL2026

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

Amanda Myntti, Jenna Kanerva, Veronika Laippala +1

In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding models on five MTEB tasks sp…

cs.CL2026

Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs

Joonatan Laato, Veera Schroderus, Jenna Kanerva +3

Digitized historical archives make it possible to study everyday social life on a large scale, but the information extracted directly from text often does not directly allow one to…

cs.CV2025

Creating a Historical Migration Dataset from Finnish Church Records, 1800-1920

Ari Vesalainen, Jenna Kanerva, Aida Nitsch +4

This article presents a large-scale effort to create a structured dataset of internal migration in Finland between 1800 and 1920 using digitized church moving records. These record…

cs.CL2025

Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs

Joonatan Laato, Jenna Kanerva, John Loehr +2

We performed a zero-shot information extraction study on a historical collection of 89,339 brief Finnish-language interviews of refugee families relocated post-WWII from Finnish Ea…

cs.CL2025

OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches

Jenna Kanerva, Cassandra Ledins, Siiri Käpyaho +1

Optical Character Recognition (OCR) systems often introduce errors when transcribing historical documents, leaving room for post-correction to improve text quality. This study eval…