collaborators

10 papers

math.OC2026

High-dimensional Limit of SGD for Diagonal Linear Networks

Begoña García Malaxechebarría, Courtney Paquette, Maryam Fazel +1

Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet…

stat.ML2026

Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models

Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1

We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction.…

math.OC2026

A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity

Damek Davis, Dmitriy Drusvyatskiy

Davis, Drusvyatskiy, and Jiang showed that gradient descent with an adaptive stepsize converges locally at a nearly-linear rate for smooth functions that grow at least quartically…

cs.LG2026

When do spectral gradient updates help in deep learning?

Damek Davis, Dmitriy Drusvyatskiy

Spectral gradient methods, such as the recently popularized Muon optimizer, are a promising alternative to standard Euclidean gradient descent for training deep neural networks and…

math.ST2026

The radius of statistical efficiency

Joshua Cutler, Mateo Díaz, Dmitriy Drusvyatskiy

Classical results in asymptotic statistics show that the Fisher information matrix controls the difficulty of estimating a statistical model from observed data. In this work, we in…

math.OC2025

Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth

Damek Davis, Dmitriy Drusvyatskiy, Liwei Jiang

A prevalent belief among optimization specialists is that linear convergence of gradient descent is contingent on the function growing quadratically away from its minimizers. In th…