5 papers
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
Katie Everett, Elliot Paquette
Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-tailed data distributions and m…
Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization
Andy Zeyi Liu, Elliot Paquette, John Sous
Training loss and throughput can hide distinct internal representation in language-model training. To examine these hidden mechanics, we use spectral measurements as practical and…
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
Ke Liang Xiao, Noah Marshall, Atish Agarwala +1
In recent years, signSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consen…
Power-Law Spectrum of the Random Feature Model
Elliot Paquette, Ke Liang Xiao, Yizhe Zhu
Scaling laws for neural networks, in which the loss decays as a power-law in the number of parameters, data, and compute, depend fundamentally on the spectral structure of the data…
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
Noah Marshall, Ke Liang Xiao, Atish Agarwala +1
The success of modern machine learning is due in part to the adaptive optimization methods that have been developed to deal with the difficulties of training large models over comp…