Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
Ke Liang Xiao, Noah Marshall, Atish Agarwala +1
In recent years, signSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consen…
stat.ML2024
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
Daniel Beaglehole, Ioannis Mitliagkas, Atish Agarwala
Understanding the mechanisms through which neural networks extract statistics from input-label pairs through feature learning is one of the most important unsolved problems in supe…
stat.ML2024
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
Noah Marshall, Ke Liang Xiao, Atish Agarwala +1
The success of modern machine learning is due in part to the adaptive optimization methods that have been developed to deal with the difficulties of training large models over comp…