2 citations · 7 across the 13 of their papers we have counts for
14 papers · 1 filter
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
Paul Jeha, Anastasiia Sedova, Louis Béthune +4
For most languages of the world, language model pre-training operates in a data-constrained regime where models must repeat their training data many times, degrading generalization…
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
Eleonora Gualdoni, Sonia Laguna, Louis Bethune +3
Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains, such as general knowledge, i…
Scaling Categorical Flow Maps
Oscar Davis, Anastasiia Filippova, Pierre Ablin +4
Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages c…
The Design Space of Tri-Modal Masked Diffusion Models
Louis Bethune, Victor Turrisi, Bruno Kacper Mlodozeniec +21
Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent work initializing and fine-tuning a base unimodal model for bimodal gen…
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
Bruno Mlodozeniec, Pierre Ablin, Louis Béthune +4
Hyperparameter tuning can dramatically impact training stability and final performance of large-scale models. Recent works on neural network parameterisations, such as P, have e…
Learning Unmasking Policies for Diffusion Language Models
Metod Jazbec, Theo X. Olausson, Louis Béthune +6
Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient…