Global optimality conditions for deep neural networks
arXiv:1707.02444
Abstract
We study the error landscape of deep linear and nonlinear neural networks with the squared error loss. Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete. For deep linear networks, we present necessary and sufficient conditions for a critical point of the risk function to be a global minimum. Surprisingly, our conditions provide an efficiently checkable test for global optimality, while such tests are typically intractable in nonconvex optimization. We further extend these results to deep nonlinear neural networks and prove similar sufficient conditions for global optimality, albeit in a more limited function space setting.
14 pages. A camera-ready version that will appear at ICLR 2018
References in corpus (2)
Cited by in corpus (28)
- Visualizing the Loss Landscape of Neural Nets
- ResNet with one-neuron hidden layers is a Universal Approximator
- Optimization for deep learning: theory and algorithms
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Understanding the Loss Surface of Neural Networks for Binary Classification
- Understanding Batch Normalization
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- Characterization of Gradient Dominance and Regularity Conditions for Neural Networks
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- On Connected Sublevel Sets in Deep Learning
- Analytic Network Learning
- Neural Architecture Search by Estimation of Network Structure Distributions
- Noether: The More Things Change, the More Stay the Same
- The Global Optimization Geometry of Shallow Linear Neural Networks
- Pure and Spurious Critical Points: a Geometric Study of Linear Networks
- Semi-Implicit Generative Model
- Measure, Manifold, Learning, and Optimization: A Theory Of Neural Networks
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- Fuzzy Logic Interpretation of Quadratic Networks
- Analysis and Optimisation of Bellman Residual Errors with Neural Function Approximation
- The Landscape of Multi-Layer Linear Neural Network From the Perspective of Algebraic Geometry
- Training Linear Neural Networks: Non-Local Convergence and Complexity Results
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum