4 papers
Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
Hugo Koubbi, Louis Hernandez, Matthieu Boussard
Low-Rank Adaptation (LoRA) is the dominant parameter-efficient fine-tuning method due to its favorable compute-performance trade-off, yet it suffers from catastrophic forgetting. W…
Homogenized Transformers
Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet
We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth…
Learning single-index models via harmonic decomposition
Nirmit Joshi, Hugo Koubbi, Theodor Misiakiewicz +1
We study the problem of learning single-index models, where the label depends on the input only through an unknown one-dimensio…
Dynamic metastability in the self-attention model
Borjan Geshkovski, Hugo Koubbi, Yury Polyanskiy +1
We consider the self-attention model - an interacting particle system on the unit sphere, which serves as a toy model for Transformers, the deep neural network architecture behind…