5 papers
High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model
O. Duranthon, F. Boncoraglio, L. Zdeborová
We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning. We introduce a sol…
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
Gabriele Farné, Fabrizio Boncoraglio, Lenka Zdeborová
A key capability of modern neural networks is their capacity to simultaneously learn underlying rules and memorize specific facts or exceptions. Yet, theoretical understanding of t…
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
Fabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani +3
Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the…
Bayes optimal learning of attention-indexed models
Fabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba +2
We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-leve…
Inference in Spreading Processes with Neural-Network Priors
Davide Ghio, Fabrizio Boncoraglio, Lenka Zdeborová
Stochastic processes on graphs are a powerful tool for modelling complex dynamical systems such as epidemics. A recent line of work focused on the inference problem where one aims…