17 citations · 70 across the 25 of their papers we have counts for
24 papers · 1 filter
Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model
Arie Wortsman-Zurich, Hugo Tabanelli, Yatin Dandi +2
We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical target that is a globally high-de…
Optimal scaling laws in learning hierarchical multi-index models
Leonardo Defilippis, Florent Krzakala, Bruno Loureiro +1
In this work, we provide a sharp theory of scaling laws for two-layer neural networks trained on a class of hierarchical multi-index targets, in a genuinely representation-limited…
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
Luca Arnaboldi, Bruno Loureiro, Ludovic Stephan +2
We study the dynamics of stochastic gradient descent (SGD) for a class of sequence models termed Sequence Single-Index (SSI) models, where the target depends on a single direction…
A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities
Yatin Dandi, Luca Pesce, Hugo Cui +3
A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to gen…
Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs
Luca Arnaboldi, Yatin Dandi, Florent Krzakala +3
We study the impact of the batch size on the iteration time of training two-layer neural networks with one-pass stochastic gradient descent (SGD) on multi-index target fu…
Dimension-free deterministic equivalents and scaling laws for random feature regression
Leonardo Defilippis, Bruno Loureiro, Theodor Misiakiewicz
In this work we investigate the generalization performance of random feature ridge regression (RFRR). Our main contribution is a general deterministic equivalent for the test error…