5 papers
High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model
O. Duranthon, F. Boncoraglio, L. Zdeborová
We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning. We introduce a sol…
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
Gabriele Farné, Fabrizio Boncoraglio, Lenka Zdeborová
A key capability of modern neural networks is their capacity to simultaneously learn underlying rules and memorize specific facts or exceptions. Yet, theoretical understanding of t…
Inference in Spreading Processes with Neural-Network Priors
Davide Ghio, Fabrizio Boncoraglio, Lenka Zdeborová
Stochastic processes on graphs are a powerful tool for modelling complex dynamical systems such as epidemics. A recent line of work focused on the inference problem where one aims…
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
Fabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani +3
Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the…
Bayes optimal learning of attention-indexed models
Fabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba +1
We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-leve…