activity
20242026
most citedAutomatic register identification for the open web using multilingual deep learning

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2026

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

Amanda Myntti, Jenna Kanerva, Veronika Laippala +1

In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding models on five MTEB tasks sp…

cs.CL20261 cited

Automatic register identification for the open web using multilingual deep learning

Erik Henriksson, Amanda Myntti, Saara Hellström +3

This article presents multilingual deep learning models for identifying web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introd…

cs.CL2025

Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation

Amanda Myntti, Erik Henriksson, Veronika Laippala +1

Pretraining data curation is a cornerstone in Large Language Model (LLM) development, leading to growing research on quality filtering of large web corpora. From statistical qualit…

cs.CL2025

An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)

Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32

Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…

cs.CL2024

Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code

Taishi Nakamura, Mayank Mishra, Simone Tedeschi +42

Pretrained language models are an integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim…