7 papers
HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
Stephan Oepen, Nikolay Arefev, Mikko Aulamo +29
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely th…
Dual-objective Language Models: Training Efficiency Without Overfitting
David Samuel, Lucas Georges Gabriel Charpentier
This paper combines autoregressive and masked-diffusion training objectives without any architectural modifications, resulting in flexible language models that outperform single-ob…
Stronger Re-identification Attacks through Reasoning and Aggregation
Lucas Georges Gabriel Charpentier, Pierre Lison
Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentione…
Systematic Generalization in Language Models Scales with Information Entropy
Sondre Wold, Lucas Georges Gabriel Charpentier, Ãtienne Simon
Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle wi…
Re-identification of De-identified Documents with Autoregressive Infilling
Lucas Georges Gabriel Charpentier, Pierre Lison
Documents revealing sensitive information about individuals must typically be de-identified. This de-identification is often done by masking all mentions of personally identifiable…
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
David Samuel, Vladislav Mikhailov, Erik Velldal +4
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages l…