3 papers
cs.LG2026
Optimal Learning Rate Schedule for Balancing Effort and Performance
Valentina Njaradi, Rodrigo Carrasco-Davis, Peter E. Latham +1
Learning how to learn efficiently is a fundamental challenge for biological agents and a growing concern for artificial ones. To learn effectively, an agent must regulate its learn…
cs.LG2025
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
Yedi Zhang, Andrew Saxe, Peter E. Latham
Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across…
cs.LG2025
Training Dynamics of In-Context Learning in Linear Attention
Yedi Zhang, Aaditya K. Singh, Peter E. Latham +1
While attention-based models have demonstrated the remarkable ability of in-context learning (ICL), the theoretical understanding of how these models acquired this ability through…