A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics
arXiv:1904.04326 · doi:10.1007/s11425-019-1628-5
Abstract
A fairly comprehensive analysis is presented for the gradient descent dynamics for training two-layer neural network models in the situation when the parameters in both layers are updated. General initialization schemes as well as general regimes for the network width and training data size are considered. In the over-parametrized regime, it is shown that gradient descent dynamics can achieve zero training loss exponentially fast regardless of the quality of the labels. In addition, it is proved that throughout the training process the functions represented by the neural network model are uniformly close to that of a kernel method. For general values of the network width and training data size, sharp estimates of the generalization error is established for target functions in the appropriate reproducing kernel Hilbert space.
Published version
References in corpus (7)
- Understanding deep learning requires rethinking generalization
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Diverse Neural Network Learns True Target Functions
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
Cited by in corpus (55)
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- A Theoretical Analysis of Deep Q-Learning
- Exploring Deep Neural Networks via Layer-Peeled Model: Minority Collapse in Imbalanced Training
- Searching the solution landscape by generalized high-index saddle dynamics
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- Two-Layer Neural Networks for Partial Differential Equations: Optimization and Generalization Theory
- Full error analysis for the training of deep neural networks
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Gradient Dynamics of Shallow Univariate ReLU Networks
- Non-convergence of stochastic gradient descent in the training of deep neural networks
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- Convergence analysis for gradient flows in the training of artificial neural networks with ReLU activation
- Linear Frequency Principle Model to Understand the Absence of Overfitting in Neural Networks
- On the Banach spaces associated with multi-layer ReLU networks: Function representation, approximation theory and gradient descent dynamics
- Neural Networks Learning and Memorization with (almost) no Over-Parameterization
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime
- Can Shallow Neural Networks Beat the Curse of Dimensionality? A mean field training perspective
- Phase diagram for two-layer ReLU neural networks at infinite-width limit
- On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
- Decoupling Gating from Linearity
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- A Priori Analysis of Stable Neural Network Solutions to Numerical PDEs
- Learning Parities with Neural Networks
- The Local Elasticity of Neural Networks
- Towards Understanding the Condensation of Neural Networks at Initial Training
- Hardness of Learning Neural Networks with Natural Weights
- Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks
- Existence, uniqueness, and convergence rates for gradient flows in the training of artificial neural networks with ReLU activation
- SPADE4: Sparsity and Delay Embedding based Forecasting of Epidemics
- Reinforcement Learning with Function Approximation: From Linear to Nonlinear
- Kolmogorov Width Decay and Poor Approximators in Machine Learning: Shallow Neural Networks, Random Feature Models and Neural Tangent Kernels
- The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models
- High-dimensional approximation spaces of artificial neural networks and applications to partial differential equations
- Generalization Guarantees for Neural Architecture Search with Train-Validation Split
- The Slow Deterioration of the Generalization Error of the Random Feature Model
- Towards an Understanding of Residual Networks Using Neural Tangent Hierarchy (NTH)
- Gradient descent provably escapes saddle points in the training of shallow ReLU networks
- Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
- Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
- A proof of convergence for the gradient descent optimization method with random initializations in the training of neural networks with ReLU activation for piecewise linear target functions
- Embedding Principle of Loss Landscape of Deep Neural Networks
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
- Strong overall error analysis for the training of artificial neural networks via random initializations
- Deep Learning without Global Optimization by Random Fourier Neural Networks
- Implicit Bias of Linear RNNs
- Learning Boolean Circuits with Neural Networks
- Dimension Independent Generalization Error by Stochastic Gradient Descent
- A priori generalization error for two-layer ReLU neural network through minimum norm solution
- A Revision of Neural Tangent Kernel-based Approaches for Neural Networks
- On the convergence of gradient descent for two layer neural networks
- Understanding the Initial Condensation of Convolutional Neural Networks