3 papers
cs.CL2026
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approxi…
cs.LG2026
Output Embedding Centering for Stable LLM Pretraining
Felix Stollenwerk, Anna Lokrantz, Niclas Hertzberg
Pretraining of large language models is not only expensive but also prone to certain training instabilities. A specific instability that often occurs at the end of training is outp…
cs.FL2025
CASP: An evaluation dataset for formal verification of C code
Niclas Hertzberg, Merlijn Sevenhuijsen, Liv KÃ¥reborn +1
Recent developments in Large Language Models (LLMs) have shown promise in automating code generation, yet the generated programs lack rigorous correctness guarantees. Formal verifi…