Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
arXiv:1902.06015
Abstract
We consider learning two layer neural networks using stochastic gradient descent. The mean-field description of this learning dynamics approximates the evolution of the network weights by an evolution in the space of probability distributions in (where is the number of parameters associated to each neuron). This evolution can be defined through a partial differential equation or, equivalently, as the gradient flow in the Wasserstein space of probability distributions. Earlier work shows that (under some regularity assumptions), the mean field description is accurate as soon as the number of hidden units is much larger than the dimension . In this paper we establish stronger and more general approximation guarantees. First of all, we show that the number of hidden units only needs to be larger than a quantity dependent on the regularity properties of the data, and independent of the dimensions. Next, we generalize this analysis to the case of unbounded activation functions, which was not covered by earlier bounds. We extend our results to noisy stochastic gradient descent. Finally, we show that kernel ridge regression can be recovered as a special limit of the mean field analysis.
61 pages
References in corpus (1)
Cited by in corpus (30)
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- A Selective Overview of Deep Learning
- A mean-field limit for certain deep neural networks
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Mathematical Models of Overparameterized Neural Networks
- Directional convergence and alignment in deep learning
- Quantitative Propagation of Chaos for SGD in Wide Neural Networks
- Functional renormalization group for multilinear disordered Langevin dynamics I: Formalism and first numerical investigations at equilibrium
- The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
- Gradients as Features for Deep Representation Learning
- Mean-field Langevin System, Optimal Control and Deep Neural Networks
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification
- The asymptotic spectrum of the Hessian of DNN throughout training
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
- Global Convergence of Three-layer Neural Networks in the Mean Field Regime
- Learning time-scales in two-layers neural networks
- A Recipe for Global Convergence Guarantee in Deep Neural Networks
- The Discovery of Dynamics via Linear Multistep Methods and Deep Learning: Error Estimation
- The sharp, the flat and the shallow: Can weakly interacting agents learn to escape bad minima?
- The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization
- Law of large numbers and central limit theorem for wide two-layer neural networks: the mini-batch and noisy case
- A Chain Graph Interpretation of Real-World Neural Networks
- Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
- A Mean-Field Theory for Learning the Schönberg Measure of Radial Basis Functions
- Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks
- One-pass Stochastic Gradient Descent in Overparametrized Two-layer Neural Networks
- Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
- Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference
- Decomposed resolution of finite-state aggregative optimal control problems
- A note on regularised NTK dynamics with an application to PAC-Bayesian training