collaborators

8 papers

stat.ML2025

The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent

Yatin Dandi, Luca Pesce, Lenka Zdeborová +2

Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce…

cs.LG2024

Asymptotic generalization error of a single-layer graph convolutional network

O. Duranthon, L. Zdeborová

While graph convolutional networks show great practical promises, the theoretical understanding of their generalization properties as a function of the number of samples is still i…

stat.ML2024

Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression

Lucas Clarté, Adrien Vandenbroucque, Guillaume Dalle +3

We investigate popular resampling methods for estimating the uncertainty of statistical models, such as subsampling, bootstrap and the jackknife, and their performance in high-dime…

cs.LG2024

A phase transition between positional and semantic learning in a solvable model of dot-product attention

Hugo Cui, Freya Behrens, Florent Krzakala +2

Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of t…

cond-mat.dis-nn2024

Quenches in the Sherrington-Kirkpatrick model

Vittorio Erba, Freya Behrens, Florent Krzakala +1

The Sherrington-Kirkpatrick (SK) model is a prototype of a complex non-convex energy landscape. Dynamical processes evolving on such landscapes and locally aiming to reach minima a…

stat.ML2024

The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents

Yatin Dandi, Emanuele Troiani, Luca Arnaboldi +3

We investigate the training dynamics of two-layer neural networks when learning multi-index target functions. We focus on multi-pass gradient descent (GD) that reuses the batches m…