3 papers
cs.LG2026
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1
In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer d…
cs.LG2025
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
Yedi Zhang, Andrew Saxe, Peter E. Latham
Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across…
cs.LG2025
Training Dynamics of In-Context Learning in Linear Attention
Yedi Zhang, Aaditya K. Singh, Peter E. Latham +1
While attention-based models have demonstrated the remarkable ability of in-context learning (ICL), the theoretical understanding of how these models acquired this ability through…