How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
arXiv:1911.12360
Abstract
A recent line of research on deep learning focuses on the extremely over-parameterized setting, and shows that when the network width is larger than a high degree polynomial of the training sample size and the inverse of the target error , deep neural networks learned by (stochastic) gradient descent enjoy nice optimization and generalization guarantees. Very recently, it is shown that under certain margin assumptions on the training data, a polylogarithmic width condition suffices for two-layer ReLU networks to converge and generalize (Ji and Telgarsky, 2019). However, whether deep neural networks can be learned with such a mild over-parameterization is still an open question. In this work, we answer this question affirmatively and establish sharper learning guarantees for deep ReLU networks trained by (stochastic) gradient descent. In specific, under certain assumptions made in previous work, our optimization and generalization guarantees hold with network width polylogarithmic in and . Our results push the study of over-parameterized deep neural networks towards more practical settings.
21 pages, 1 figure, 1 table. In ICLR 2021
References in corpus (9)
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Norm-Based Capacity Control in Neural Networks
- Diverse Neural Network Learns True Target Functions
- Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- Algorithm-Dependent Generalization Bounds for Overparameterized Deep Residual Networks
Cited by in corpus (22)
- Towards Understanding the Spectral Bias of Deep Learning
- Two-Layer Neural Networks for Partial Differential Equations: Optimization and Generalization Theory
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
- Recent advances in deep learning theory
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- A Corrective View of Neural Networks: Representation, Memorization and Learning
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear Widths
- Reproducing Activation Function for Deep Learning
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
- The Discovery of Dynamics via Linear Multistep Methods and Deep Learning: Error Estimation
- Is deeper better? It depends on locality of relevant features
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
- Training Two-Layer ReLU Networks with Gradient Descent is Inconsistent
- The Dynamics of Gradient Descent for Overparametrized Neural Networks
- Label-Aware Neural Tangent Kernel: Toward Better Generalization and Local Elasticity
- Actor-critic is implicitly biased towards high entropy optimal policies
- Achieving Small Test Error in Mildly Overparameterized Neural Networks
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- On the Provable Generalization of Recurrent Neural Networks
- On the Stability Properties and the Optimization Landscape of Training Problems with Squared Loss for Neural Networks and General Nonlinear Conic Approximation Schemes
- On Assessing the Quantum Advantage for MaxCut Provided by Quantum Neural Network Ansätze
- Training Linear Neural Networks: Non-Local Convergence and Complexity Results