4 papers
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…
Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke
Yu Wu, Ananth Mahadevan, Filip Ginter +2
While digitized corpora have transformed the study of intellectual transmission, current methods rely heavily on lexical text reuse detection, capturing verbatim quotations but fun…
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering
Erik Henriksson, Otto Tarkka, Filip Ginter
Data quality is crucial for training Large Language Models (LLMs). Traditional heuristic filters often miss low-quality text or mistakenly remove valuable content. In this paper, w…
Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
Emil Nuutinen, Iiro Rastas, Filip Ginter
We apply a simple method to machine translate datasets with span-level annotation using the DeepL MT service and its ability to translate formatted documents. Using this method, we…