A mean-field limit for certain deep neural networks
arXiv:1906.00193
Abstract
Understanding deep neural networks (DNNs) is a key challenge in the theory of machine learning, with potential applications to the many fields where DNNs have been successfully used. This article presents a scaling limit for a DNN being trained by stochastic gradient descent. Our networks have a fixed (but arbitrary) number of inner layers; neurons per layer; full connections between layers; and fixed weights (or "random features" that are not trained) near the input and output. Our results describe the evolution of the DNN during training in the limit when , which we relate to a mean field model of McKean-Vlasov type. Specifically, we show that network weights are approximated by certain "ideal particles" whose distribution and dependencies are described by the mean-field model. A key part of the proof is to show existence and uniqueness for our McKean-Vlasov problem, which does not seem to be amenable to existing theory. Our paper extends previous work on the case by Mei, Montanari and Nguyen; Rotskoff and Vanden-Eijnden; and Sirignano and Spiliopoulos. We also complement recent independent work on by Sirignano and Spiliopoulos (who consider a less natural scaling limit) and Nguyen (who nonrigorously derives similar results).
79 pages and 2 figures
References in corpus (3)
Cited by in corpus (8)
- Optimization for deep learning: theory and algorithms
- Mathematical Models of Overparameterized Neural Networks
- On the Banach spaces associated with multi-layer ReLU networks: Function representation, approximation theory and gradient descent dynamics
- Can Shallow Neural Networks Beat the Curse of Dimensionality? A mean field training perspective
- On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- A Note on the Global Convergence of Multilayer Neural Networks in the Mean Field Regime
- Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks