Nearly Minimal Over-Parametrization of Shallow Neural Networks
arXiv:1910.03948
Abstract
A recent line of work has shown that an overparametrized neural network can perfectly fit the training data, an otherwise often intractable nonconvex optimization problem. For (fully-connected) shallow networks, in the best case scenario, the existing theory requires quadratic over-parametrization as a function of the number of training samples. This paper establishes that linear overparametrization is sufficient to fit the training data, using a simple variant of the (stochastic) gradient descent. Crucially, unlike several related works, the training considered in this paper is not limited to the lazy regime in the sense cautioned against in [1, 2]. Beyond shallow networks, the framework developed in this work for over-parametrization is applicable to a variety of learning problems.
This paper is submitted without consent of the co-authors
References in corpus (6)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- An Improved Analysis of Training Over-parameterized Deep Neural Networks
- Learning Non-overlapping Convolutional Neural Networks with Multiple Kernels