On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
arXiv:2005.13530
Abstract
We describe a necessary and sufficient condition for the convergence to minimum Bayes risk when training two-layer ReLU-networks by gradient descent in the mean field regime with omni-directional initial parameter distribution. This article extends recent results of Chizat and Bach to ReLU-activated networks and to the situation in which there are no parameters which exactly achieve MBR. The condition does not depend on the initalization of parameters and concerns only the weak convergence of the realization of the neural network, not its parameter distribution.
References in corpus (6)
- A Priori Estimates of the Population Risk for Residual Networks
- A mean-field limit for certain deep neural networks
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- Can Shallow Neural Networks Beat the Curse of Dimensionality? A mean field training perspective
- Kolmogorov Width Decay and Poor Approximators in Machine Learning: Shallow Neural Networks, Random Feature Models and Neural Tangent Kernels