4 papers · 1 filter
High-dimensional Limit of SGD for Diagonal Linear Networks
Begoña GarcÃa MalaxechebarrÃa, Courtney Paquette, Maryam Fazel +1
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet…
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
Damek Davis, Dmitriy Drusvyatskiy
Davis, Drusvyatskiy, and Jiang showed that gradient descent with an adaptive stepsize converges locally at a nearly-linear rate for smooth functions that grow at least quartically…
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
Damek Davis, Dmitriy Drusvyatskiy, Liwei Jiang
A prevalent belief among optimization specialists is that linear convergence of gradient descent is contingent on the function growing quadratically away from its minimizers. In th…
Invariant Kernels: Rank Stabilization and Generalization Across Dimensions
Mateo DÃaz, Dmitriy Drusvyatskiy, Jack Kendrick +1
Symmetry arises often when learning from high dimensional data. For example, data sets consisting of point clouds, graphs, and unordered sets appear routinely in contemporary appli…