activity
20242026
collaborators
Showing cs.LGShow all

12 papers · 1 filter

cs.LG2026

Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining

Thiziri Nait Saada, Louis Bethune, Michal Klein +3

Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quali…

cs.LG2026

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing

Anastasiia Filippova, David Grangier, Marco Cuturi +1

Serving transformer language models with high throughput requires caching Key-Values (KVs) to avoid redundant computation during autoregressive generation. The memory footprint of…

cs.LG2026

Compute-Optimal Quantization-Aware Training

Aleksandr Dremov, David Grangier, Angelos Katharopoulos +1

Quantization-aware training (QAT) is a leading technique for improving the accuracy of quantized neural networks. Previous work has shown that decomposing training into a full-prec…

cs.LG2025

Scaling Laws for Optimal Data Mixtures

Mustafa Shukor, Louis Bethune, Dan Busbridge +4

Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance…

cs.LG2025

Partial Parameter Updates for Efficient Distributed Training

Anastasiia Filippova, Angelos Katharopoulos, David Grangier +1

We introduce a memory- and compute-efficient method for low-communication distributed training. Existing methods reduce communication by performing multiple local updates between i…

cs.LG2025

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Louis Bethune, David Grangier, Dan Busbridge +3

A widespread strategy to obtain a language model that performs well on a target domain is to finetune a pretrained model to perform unsupervised next-token prediction on data from…