Can Shallow Neural Networks Beat the Curse of Dimensionality? A mean field training perspective
arXiv:2005.10815
Abstract
We prove that the gradient descent training of a two-layer neural network on empirical or population risk may not decrease population risk at an order faster than under mean field scaling. Thus gradient descent training for fitting reasonably smooth, but truly high-dimensional data may be subject to the curse of dimensionality. We present numerical evidence that gradient descent training with general Lipschitz target functions becomes slower and slower as the dimension increases, but converges at approximately the same rate in all dimensions when the target function lies in the natural function space for two-layer ReLU networks.
5 figures
References in corpus (3)
Cited by in corpus (7)
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- On the Banach spaces associated with multi-layer ReLU networks: Function representation, approximation theory and gradient descent dynamics
- On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
- Kolmogorov Width Decay and Poor Approximators in Machine Learning: Shallow Neural Networks, Random Feature Models and Neural Tangent Kernels
- The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models
- Sinc Kolmogorov-Arnold network and its application for solving PDEs with singularities
- An invariance principle for gradient flows in the space of probability measures