5 papers
A Defense of the Quadratic Model
Alexandru Meterez, Pranav Ajit Nair, Depen Morwani +3
Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically trac…
Improved high-dimensional estimation with Langevin dynamics and stochastic weight averaging
Stanley Wei, Alex Damian, Jason D. Lee
Significant recent work has studied the ability of gradient descent to recover a hidden planted direction in different high-dimensional settings, including t…
Understanding Optimization in Deep Learning with Central Flows
Jeremy M. Cohen, Alex Damian, Ameet Talwalkar +2
Traditional theories of optimization cannot describe the dynamics of optimization in deep learning, even in the simple setting of deterministic training. The challenge is that opti…
The Generative Leap: Sharp Sample Complexity for Efficiently Learning Gaussian Multi-Index Models
Alex Damian, Jason D. Lee, Joan Bruna
In this work we consider generic Gaussian Multi-index models, in which the labels only depend on the (Gaussian) -dimensional inputs through their projection onto a low-dimension…
Learning Compositional Functions with Transformers from Easy-to-Hard Data
Zixuan Wang, Eshaan Nichani, Alberto Bietti +4
Transformer-based language models have demonstrated impressive capabilities across a range of complex reasoning tasks. Prior theoretical work exploring the expressive power of tran…