activity
20242026
collaborators

7 papers

cs.CL2026

HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

Stephan Oepen, Nikolay Arefev, Mikko Aulamo +29

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely th…

cs.CL2026

Dual-objective Language Models: Training Efficiency Without Overfitting

David Samuel, Lucas Georges Gabriel Charpentier

This paper combines autoregressive and masked-diffusion training objectives without any architectural modifications, resulting in flexible language models that outperform single-ob…

cs.CL2025

Stronger Re-identification Attacks through Reasoning and Aggregation

Lucas Georges Gabriel Charpentier, Pierre Lison

Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentione…

cs.CL2025

Systematic Generalization in Language Models Scales with Information Entropy

Sondre Wold, Lucas Georges Gabriel Charpentier, Étienne Simon

Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle wi…

cs.CL2025

Re-identification of De-identified Documents with Autoregressive Infilling

Lucas Georges Gabriel Charpentier, Pierre Lison

Documents revealing sensitive information about individuals must typically be de-identified. This de-identification is often done by masking all mentions of personally identifiable…

cs.CL2025

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

David Samuel, Vladislav Mikhailov, Erik Velldal +4

Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages l…