6 papers
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
Amanda Myntti, Jenna Kanerva, Veronika Laippala +1
In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding models on five MTEB tasks sp…
Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs
Joonatan Laato, Veera Schroderus, Jenna Kanerva +3
Digitized historical archives make it possible to study everyday social life on a large scale, but the information extracted directly from text often does not directly allow one to…
Creating a Historical Migration Dataset from Finnish Church Records, 1800-1920
Ari Vesalainen, Jenna Kanerva, Aida Nitsch +4
This article presents a large-scale effort to create a structured dataset of internal migration in Finland between 1800 and 1920 using digitized church moving records. These record…
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs
Joonatan Laato, Jenna Kanerva, John Loehr +2
We performed a zero-shot information extraction study on a historical collection of 89,339 brief Finnish-language interviews of refugee families relocated post-WWII from Finnish Ea…
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
Jenna Kanerva, Cassandra Ledins, Siiri Käpyaho +1
Optical Character Recognition (OCR) systems often introduce errors when transcribing historical documents, leaving room for post-correction to improve text quality. This study eval…