10 papers
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
Yatin Dandi, Luca Pesce, Lenka Zdeborová +1
Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce…
Quenches in the Sherrington-Kirkpatrick model
Vittorio Erba, Freya Behrens, Florent Krzakala +1
The Sherrington-Kirkpatrick (SK) model is a prototype of a complex non-convex energy landscape. Dynamical processes evolving on such landscapes and locally aiming to reach minima a…
Fundamental limits of Non-Linear Low-Rank Matrix Estimation
Pierre Mergny, Justin Ko, Florent Krzakala +1
We consider the task of estimating a low-rank matrix from non-linear and noisy observations. We prove a strong universality result showing that Bayes-optimal performances are chara…
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
Lucas Clarté, Adrien Vandenbroucque, Guillaume Dalle +3
We investigate popular resampling methods for estimating the uncertainty of statistical models, such as subsampling, bootstrap and the jackknife, and their performance in high-dime…
Asymptotics of feature learning in two-layer networks after one gradient-step
Hugo Cui, Luca Pesce, Yatin Dandi +4
In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single grad…
A phase transition between positional and semantic learning in a solvable model of dot-product attention
Hugo Cui, Freya Behrens, Florent Krzakala +1
Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of t…