8 papers
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
Yatin Dandi, Luca Pesce, Lenka Zdeborová +2
Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce…
Asymptotic generalization error of a single-layer graph convolutional network
O. Duranthon, L. Zdeborová
While graph convolutional networks show great practical promises, the theoretical understanding of their generalization properties as a function of the number of samples is still i…
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
Lucas Clarté, Adrien Vandenbroucque, Guillaume Dalle +3
We investigate popular resampling methods for estimating the uncertainty of statistical models, such as subsampling, bootstrap and the jackknife, and their performance in high-dime…
A phase transition between positional and semantic learning in a solvable model of dot-product attention
Hugo Cui, Freya Behrens, Florent Krzakala +2
Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of t…
Quenches in the Sherrington-Kirkpatrick model
Vittorio Erba, Freya Behrens, Florent Krzakala +1
The Sherrington-Kirkpatrick (SK) model is a prototype of a complex non-convex energy landscape. Dynamical processes evolving on such landscapes and locally aiming to reach minima a…
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
Yatin Dandi, Emanuele Troiani, Luca Arnaboldi +3
We investigate the training dynamics of two-layer neural networks when learning multi-index target functions. We focus on multi-pass gradient descent (GD) that reuses the batches m…