Are ResNets Provably Better than Linear Predictors?
arXiv:1804.06739
Abstract
A residual network (or ResNet) is a standard deep neural net architecture, with state-of-the-art performance across numerous applications. The main premise of ResNets is that they allow the training of each layer to focus on fitting just the residual of the previous layer's output and the target output. Thus, we should expect that the trained network is no worse than what we can obtain if we remove the residual layers and train a shallower network instead. However, due to the non-convexity of the optimization problem, it is not at all clear that ResNets indeed achieve this behavior, rather than getting stuck at some arbitrarily poor local minimum. In this paper, we rigorously prove that arbitrarily deep, nonlinear residual units indeed exhibit this behavior, in the sense that the optimization landscape contains no local minima with value above what can be obtained with a linear predictor (namely a 1-layer network). Notably, we show this under minimal or no assumptions on the precise network architecture, data distribution, or loss function used. We also provide a quantitative analysis of approximate stationary points for this problem. Finally, we show that with a certain tweak to the architecture, training the network with standard stochastic gradient descent achieves an objective value close or better than any linear predictor.
Comparison to previous arXiv version: Minor changes to incorporate comments of NIPS 2018 reviewers (main results are unaffected)
Cited by in corpus (15)
- ResNet with one-neuron hidden layers is a Universal Approximator
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- Spurious Valleys in Two-layer Neural Network Optimization Landscapes
- Effect of Depth and Width on Local Minima in Deep Learning
- Elimination of All Bad Local Minima in Deep Learning
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
- Weakly-Supervised Action Localization and Action Recognition using Global-Local Attention of 3D CNN
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers
- Geometric Insights into the Convergence of Nonlinear TD Learning
- Are deep ResNets provably better than linear predictors?
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 1: Theory
- Boosting Offline Reinforcement Learning with Residual Generative Modeling
- ResNEsts and DenseNEsts: Block-based DNN Models with Improved Representation Guarantees
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations