Mathematical Models of Overparameterized Neural Networks
arXiv:2012.13982 · doi:10.1109/JPROC.2020.3048020
Abstract
Deep learning has received considerable empirical successes in recent years. However, while many ad hoc tricks have been discovered by practitioners, until recently, there has been a lack of theoretical understanding for tricks invented in the deep learning literature. Known by practitioners that overparameterized neural networks are easy to learn, in the past few years there have been important theoretical developments in the analysis of overparameterized neural networks. In particular, it was shown that such systems behave like convex systems under various restricted settings, such as for two-layer NNs, and when learning is restricted locally in the so-called neural tangent kernel space around specialized initializations. This paper discusses some of these recent progresses leading to significant better understanding of neural networks. We will focus on the analysis of two-layer neural networks, and explain the key mathematical models, with their algorithmic implications. We will then discuss challenges in understanding deep neural networks and some current research directions.
References in corpus (15)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- How to Escape Saddle Points Efficiently
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Why Deep Neural Networks for Function Approximation?
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
- Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?
- An Improved Analysis of Training Over-parameterized Deep Neural Networks
- A mean-field limit for certain deep neural networks
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Sharp Analysis for Nonconvex SGD Escaping from Saddle Points
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- The Local Elasticity of Neural Networks
- Convex Formulation of Overparameterized Deep Neural Networks