collaborators

6 papers

cs.LG2026

How are linear representations learned? Exact solutions to the dynamics of abstraction

William W. Yang, Andrew M. Saxe, Peter E. Latham

In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear…

cs.LG2026

Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing

Valentina Njaradi, Clémentine Dominé, Rachel Swanson +2

Learning to generalise from limited data is a fundamental challenge for both artificial and biological systems. A common strategy is to extract reusable structure from abundant unl…

cs.LG2026

Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures

Yedi Zhang, Andrew Saxe, Peter E. Latham

Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across…

cs.LG2026

Optimal Learning Rate Schedule for Balancing Effort and Performance

Valentina Njaradi, Rodrigo Carrasco-Davis, Peter E. Latham +1

Learning how to learn efficiently is a fundamental challenge for biological agents and a growing concern for artificial ones. To learn effectively, an agent must regulate its learn…

cs.LG2025

Training Dynamics of In-Context Learning in Linear Attention

Yedi Zhang, Aaditya K. Singh, Peter E. Latham +1

While attention-based models have demonstrated the remarkable ability of in-context learning (ICL), the theoretical understanding of how these models acquired this ability through…

cs.LG2025

When Are Bias-Free ReLU Networks Effectively Linear Networks?

Yedi Zhang, Andrew Saxe, Peter E. Latham

We investigate the implications of removing bias in ReLU networks regarding their expressivity and learning dynamics. We first show that two-layer bias-free ReLU networks have limi…