Optimization Landscapes of Wide Deep Neural Networks Are Benign
arXiv:2010.00885
Abstract
We analyze the optimization landscapes of deep learning with wide networks. We highlight the importance of constraints for such networks and show that constraint -- as well as unconstraint -- empirical-risk minimization over such networks has no confined points, that is, suboptimal parameters that are difficult to escape from. Hence, our theories substantiate the common belief that wide neural networks are not only highly expressive but also comparably easy to optimize.
References in corpus (12)
- Nearly unbiased variable selection under minimax concave penalty
- A Convergence Theory for Deep Learning via Over-Parameterization
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Qualitatively characterizing neural network optimization problems
- On the Computational Efficiency of Training Neural Networks
- Sparse-Input Neural Networks for High-dimensional Nonparametric Regression and Classification
- Optimization Landscape and Expressivity of Deep CNNs
- Approximation and Estimation for High-Dimensional Deep Learning Networks
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- Complexity, Statistical Risk, and Metric Entropy of Deep Nets Using Total Path Variation
- Risk Bounds for Robust Deep Learning
- Non-attracting Regions of Local Minima in Deep and Wide Neural Networks