Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model
O. Duranthon, F. Boncoraglio, L. Zdeborová
We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning. We introduce a sol…
cs.LG2026
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
O. Duranthon, P. Marion, C. Boyer +2
Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly underst…
cs.LG2024
Asymptotic generalization error of a single-layer graph convolutional network
O. Duranthon, L. Zdeborová
While graph convolutional networks show great practical promises, the theoretical understanding of their generalization properties as a function of the number of samples is still i…