4 papers
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1
In this short note we consider the gradient descent dynamics of deep scalar linear networks, , which enjoy exact time-course solutions for any integer d…
A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning
Nicolas Anguita, Francesco Locatello, Andrew M. Saxe +4
Pretraining and fine-tuning are central stages in modern machine learning systems. In practice, feature learning plays an important role across both stages: deep neural networks le…
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
Loek van Rossem, Andrew M. Saxe
Even when massively overparameterized, deep neural networks show a remarkable ability to generalize. Research on this phenomenon has focused on generalization within distribution,…
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
Devon Jarvis, Richard Klein, Benjamin Rosman +1
In spite of finite dimension ReLU neural networks being a consistent factor behind recent deep learning successes, a theory of feature learning in these models remains elusive. Cur…