10 papers
High-dimensional Limit of SGD for Diagonal Linear Networks
Begoña GarcÃa MalaxechebarrÃa, Courtney Paquette, Maryam Fazel +1
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet…
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1
We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction.…
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
Damek Davis, Dmitriy Drusvyatskiy
Davis, Drusvyatskiy, and Jiang showed that gradient descent with an adaptive stepsize converges locally at a nearly-linear rate for smooth functions that grow at least quartically…
When do spectral gradient updates help in deep learning?
Damek Davis, Dmitriy Drusvyatskiy
Spectral gradient methods, such as the recently popularized Muon optimizer, are a promising alternative to standard Euclidean gradient descent for training deep neural networks and…
The radius of statistical efficiency
Joshua Cutler, Mateo DÃaz, Dmitriy Drusvyatskiy
Classical results in asymptotic statistics show that the Fisher information matrix controls the difficulty of estimating a statistical model from observed data. In this work, we in…
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
Damek Davis, Dmitriy Drusvyatskiy, Liwei Jiang
A prevalent belief among optimization specialists is that linear convergence of gradient descent is contingent on the function growing quadratically away from its minimizers. In th…