4 papers
High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model
O. Duranthon, F. Boncoraglio, L. Zdeborová
We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning. We introduce a sol…
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
O. Duranthon, P. Marion, C. Boyer +2
Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly underst…
Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model
O. Duranthon, L. Zdeborová
Graph neural networks (GNNs) are designed to process data associated with graphs. They are finding an increasing range of applications; however, as with other modern machine learni…
Asymptotic generalization error of a single-layer graph convolutional network
O. Duranthon, L. Zdeborová
While graph convolutional networks show great practical promises, the theoretical understanding of their generalization properties as a function of the number of samples is still i…