Depth Creates No Bad Local Minima
arXiv:1702.08580
Abstract
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatly simplify a recently proposed proof to show that all of the local minima of feedforward deep linear neural networks are global minima. Our theoretical results generalize previous results with fewer assumptions, and this analysis provides a method to show similar results beyond square loss in deep linear models.
References in corpus (5)
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Exponentially vanishing sub-optimal local minima in multilayer neural networks
- Local minima in training of neural networks
- Distribution-Specific Hardness of Learning Neural Networks
Cited by in corpus (45)
- Visualizing the Loss Landscape of Neural Nets
- Optimization for deep learning: theory and algorithms
- The Global Landscape of Neural Networks: An Overview
- Energy-entropy competition and the effectiveness of stochastic gradient descent in machine learning
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Global optimality conditions for deep neural networks
- Understanding Batch Normalization
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- Exponential Convergence Time of Gradient Descent for One-Dimensional Deep Linear Neural Networks
- Recent advances in deep learning theory
- Understanding Deep Learning via Decision Boundary
- Convergence to Second-Order Stationarity for Constrained Non-Convex Optimization
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Decoupling Gating from Linearity
- Depth creates no more spurious local minima
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- Visualized Insights into the Optimization Landscape of Fully Convolutional Networks
- On the Global Convergence of Training Deep Linear ResNets
- Classification Logit Two-sample Testing by Neural Networks
- Deep linear neural networks with arbitrary loss: All local minima are global
- Landscape Complexity for the Empirical Risk of Generalized Linear Models
- Analytic Network Learning
- On Connected Sublevel Sets in Deep Learning
- Non-attracting Regions of Local Minima in Deep and Wide Neural Networks
- Pure and Spurious Critical Points: a Geometric Study of Linear Networks
- The Global Optimization Geometry of Shallow Linear Neural Networks
- The global optimum of shallow neural network is attained by ridgelet transform
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 1: Theory
- Sub-Optimal Local Minima Exist for Neural Networks with Almost All Non-Linear Activations
- On the achievability of blind source separation for high-dimensional nonlinear source mixtures
- The loss landscape of deep linear neural networks: a second-order analysis
- Ridge Regression with Over-Parametrized Two-Layer Networks Converge to Ridgelet Spectrum
- Who is Afraid of Big Bad Minima? Analysis of Gradient-Flow in a Spiked Matrix-Tensor Model
- Adaptive Stochastic Gradient Langevin Dynamics: Taming Convergence and Saddle Point Escape Time
- A Unified Framework for Training Neural Networks
- The Landscape of Multi-Layer Linear Neural Network From the Perspective of Algebraic Geometry
- Interpretable Few-Shot Learning via Linear Distillation
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- Training Linear Neural Networks: Non-Local Convergence and Complexity Results
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- When Are Solutions Connected in Deep Networks?
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Notes on Deep Learning Theory