5 papers · 1 filter
Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
Ioannis Bantzis, James B. Simon, Arthur Jacot
When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter space. We study the so-called escap…
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
Enric Boix-Adsera, Neil Mallinar, James B. Simon +1
It is a central challenge in deep learning to understand how neural networks learn representations. A leading approach is the Neural Feature Ansatz (NFA) (Radhakrishnan et al. 2024…
The Optimization Landscape of SGD Across the Feature Learning Strength
Alexander Atanasov, Alexandru Meterez, James B. Simon +1
We consider neural networks (NNs) where the final layer is down-scaled by a fixed hyperparameter . Recent work has identified as controlling the strength of feature learni…
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
Neil Mallinar, James B. Simon, Amirhesam Abedsoltan +3
The practical success of overparameterized neural networks has motivated the recent scientific study of interpolating methods, which perfectly fit their training data. Certain inte…
A Spectral Condition for Feature Learning
Greg Yang, James B. Simon, Jeremy Bernstein
The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network'…