6 papers
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1
In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer d…
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
Sara DragutinoviÄ, Yedi Zhang, Rajesh Ranganath
While Adam has long been the ubiquitous default optimizer for deep neural networks, Muon has recently seen rapid adoption due to its superior training speed. Although much of the l…
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
Yedi Zhang, Andrew Saxe, Peter E. Latham
Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across…
Training Dynamics of In-Context Learning in Linear Attention
Yedi Zhang, Aaditya K. Singh, Peter E. Latham +1
While attention-based models have demonstrated the remarkable ability of in-context learning (ICL), the theoretical understanding of how these models acquired this ability through…
When Are Bias-Free ReLU Networks Effectively Linear Networks?
Yedi Zhang, Andrew Saxe, Peter E. Latham
We investigate the implications of removing bias in ReLU networks regarding their expressivity and learning dynamics. We first show that two-layer bias-free ReLU networks have limi…
Understanding Unimodal Bias in Multimodal Deep Linear Networks
Yedi Zhang, Peter E. Latham, Andrew Saxe
Using multiple input streams simultaneously to train multimodal neural networks is intuitively advantageous but practically challenging. A key challenge is unimodal bias, where a n…