6 papers · 1 filter
Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model
Arie Wortsman-Zurich, Hugo Tabanelli, Yatin Dandi +2
We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical target that is a globally high-de…
Deep Learning of Compositional Targets with Hierarchical Spectral Methods
Hugo Tabanelli, Yatin Dandi, Luca Pesce +1
Why depth yields a genuine computational advantage over shallow methods remains a central open question in learning theory. We study this question in a controlled high-dimensional…
Asymptotics of Non-Convex Generalized Linear Models in High-Dimensions: A proof of the replica formula
Matteo Vilucchio, Yatin Dandi, Matéo Pirio Rossignol +2
The analytic characterization of the high-dimensional behavior of optimization for Generalized Linear Models (GLMs) with Gaussian data has been a central focus in statistics and pr…
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
Yatin Dandi, Luca Pesce, Lenka Zdeborová +2
Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce…
A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities
Yatin Dandi, Luca Pesce, Hugo Cui +3
A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to gen…
Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs
Luca Arnaboldi, Yatin Dandi, Florent Krzakala +3
We study the impact of the batch size on the iteration time of training two-layer neural networks with one-pass stochastic gradient descent (SGD) on multi-index target fu…