14 papers
Dynamics of learning to integrate in linear recurrent neural networks
Blake Bordelon, Jordan Cotler, Cengiz Pehlevan +1
Learning recurrent connectivity that supports memory over long intrinsic timescales is a basic problem in the theory of dynamical computation. While continuous attractor and integr…
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
Clarissa Lauditi, Cengiz Pehlevan, Blake Bordelon
We study the evolution of hidden-weight spectra in wide neural networks trained by (stochastic) gradient descent. We develop a two-level dynamical mean-field theory (DMFT) that joi…
Hyperparameter Transfer with Mixture-of-Expert Layers
Tianze Jiang, Blake Bordelon, Cengiz Pehlevan +1
Mixture-of-Experts (MoE) layers have emerged as an important tool in scaling up modern neural networks by decoupling total trainable parameters from activated parameters in the for…
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
Blake Bordelon, Francesco Mori
Setting the learning rate (LR) for a deep learning model is a critical part of successful training. Choosing LRs is often done empirically with trial and error. In this work, we ex…
There Will Be a Scientific Theory of Deep Learning
Jamie Simon, Daniel Kunin, Alexander Atanasov +11
In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the train…
Transfer Learning in Infinite Width Feature Learning Networks
Clarissa Lauditi, Blake Bordelon, Cengiz Pehlevan
We develop a theory of transfer learning in infinitely wide neural networks under gradient flow that quantifies when pretraining on a source task improves generalization on a targe…