activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape

Ioannis Bantzis, James B. Simon, Arthur Jacot

When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter space. We study the so-called escap…

cs.LG2025

The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations

Enric Boix-Adsera, Neil Mallinar, James B. Simon +1

It is a central challenge in deep learning to understand how neural networks learn representations. A leading approach is the Neural Feature Ansatz (NFA) (Radhakrishnan et al. 2024…

cs.LG2025

The Optimization Landscape of SGD Across the Feature Learning Strength

Alexander Atanasov, Alexandru Meterez, James B. Simon +1

We consider neural networks (NNs) where the final layer is down-scaled by a fixed hyperparameter . Recent work has identified as controlling the strength of feature learni…

cs.LG2024

Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting

Neil Mallinar, James B. Simon, Amirhesam Abedsoltan +3

The practical success of overparameterized neural networks has motivated the recent scientific study of interpolating methods, which perfectly fit their training data. Certain inte…

cs.LG2024

A Spectral Condition for Feature Learning

Greg Yang, James B. Simon, Jeremy Bernstein

The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network'…