activity
20242026
collaborators

6 papers

cs.LG2026

Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1

In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer d…

cs.LG2026

To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters

Sara Dragutinović, Yedi Zhang, Rajesh Ranganath

While Adam has long been the ubiquitous default optimizer for deep neural networks, Muon has recently seen rapid adoption due to its superior training speed. Although much of the l…

cs.LG2026

Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures

Yedi Zhang, Andrew Saxe, Peter E. Latham

Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across…

cs.LG2025

Training Dynamics of In-Context Learning in Linear Attention

Yedi Zhang, Aaditya K. Singh, Peter E. Latham +1

While attention-based models have demonstrated the remarkable ability of in-context learning (ICL), the theoretical understanding of how these models acquired this ability through…

cs.LG2025

When Are Bias-Free ReLU Networks Effectively Linear Networks?

Yedi Zhang, Andrew Saxe, Peter E. Latham

We investigate the implications of removing bias in ReLU networks regarding their expressivity and learning dynamics. We first show that two-layer bias-free ReLU networks have limi…

cs.LG2024

Understanding Unimodal Bias in Multimodal Deep Linear Networks

Yedi Zhang, Peter E. Latham, Andrew Saxe

Using multiple input streams simultaneously to train multimodal neural networks is intuitively advantageous but practically challenging. A key challenge is unimodal bias, where a n…