Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
Ke Liang Xiao, Noah Marshall, Atish Agarwala +1
In recent years, signSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consen…
stat.ML2024
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
Noah Marshall, Ke Liang Xiao, Atish Agarwala +1
The success of modern machine learning is due in part to the adaptive optimization methods that have been developed to deal with the difficulties of training large models over comp…