1 citations · 1 across the 2 of their papers we have counts for
3 papers
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
Valérie Castin, Kimia Nadjahi, Pierre Ablin +1
Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multiple pairs of low-rank factor…
A Unified Perspective on the Dynamics of Deep Transformers
Valérie Castin, Pierre Ablin, José Antonio Carrillo +1
Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens. This representation is then exploited by the atte…
How Smooth Is Attention?
Valérie Castin, Pierre Ablin, Gabriel Peyré
Self-attention and masked self-attention are at the heart of Transformers' outstanding success. Still, our mathematical understanding of attention, in particular of its Lipschitz p…