activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Pretraining with hierarchical memories: separating long-tail and common knowledge

Hadi Pouransari, David Grangier, C Thomas +2

The impressive performance gains of modern language models currently rely on scaling parameters: larger models store more world knowledge and reason better. Yet compressing all wor…

cs.CL2026

Optimal Splitting of Language Models from Mixtures to Specialized Domains

Skyler Seto, Pierre Ablin, Anastasiia Filippova +4

Language models achieve impressive performance on a variety of knowledge, language, and reasoning tasks due to the scale and diversity of pretraining data available. The standard t…

cs.CL2025

Assessing the Role of Data Quality in Training Bilingual Language Models

Skyler Seto, Maartje ter Hoeve, Maureen de Seyssel +1

Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly betw…

cs.CL2025

Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

David Grangier, Simin Fan, Skyler Seto +1

Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain…

cs.CL2025

Training Bilingual LMs with Data Constraints in the Targeted Language

Skyler Seto, Maartje ter Hoeve, Richard He Bai +2

Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretrai…