Identity Matters in Deep Learning
arXiv:1611.04231
Abstract
An emerging design principle in deep learning is that each layer of a deep artificial neural network should be able to easily express the identity transformation. This idea not only motivated various normalization techniques, such as \emph{batch normalization}, but was also key to the immense success of \emph{residual networks}. In this work, we put the principle of \emph{identity parameterization} on a more solid theoretical footing alongside further empirical progress. We first give a strikingly simple proof that arbitrarily deep linear residual networks have no spurious local optima. The same result for linear feed-forward networks in their standard parameterization is substantially more delicate. Second, we show that residual networks with ReLu activations have universal finite-sample expressivity in the sense that the network can represent any function of its sample provided that the model has more parameters than the sample size. Directly inspired by our theory, we experiment with a radically simple residual architecture consisting of only residual convolutional layers and ReLu activations, but no batch normalization, dropout, or max pool. Our model improves significantly on previous all-convolutional networks on the CIFAR10, CIFAR100, and ImageNet classification benchmarks.
ICLR 2017; fixed minor typos in the previous version
Cited by in corpus (51)
- A Convergence Theory for Deep Learning via Over-Parameterization
- Visualizing the Loss Landscape of Neural Nets
- How Does Batch Normalization Help Optimization?
- On the Convergence Rate of Training Recurrent Neural Networks
- Are All Layers Created Equal?
- A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks
- A Deep Learning Approach for Blind Drift Calibration of Sensor Networks
- Implicit Regularization in Deep Matrix Factorization
- Multi-level Residual Networks from Dynamical Systems View
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- Depth with Nonlinearity Creates No Bad Local Minima in ResNets
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Neural Stochastic Differential Equations: Deep Latent Gaussian Models in the Diffusion Limit
- Understanding the Loss Surface of Neural Networks for Binary Classification
- Adding One Neuron Can Eliminate All Bad Local Minima
- Elimination of All Bad Local Minima in Deep Learning
- RMP-SNN: Residual Membrane Potential Neuron for Enabling Deeper High-Accuracy and Low-Latency Spiking Neural Network
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- Robust Bi-Tempered Logistic Loss Based on Bregman Divergences
- Mathematical Models of Overparameterized Neural Networks
- Algorithmic Regularization in Over-parameterized Matrix Sensing and Neural Networks with Quadratic Activations
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Representing smooth functions as compositions of near-identity functions with implications for deep network optimization
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- PolyGAN: High-Order Polynomial Generators
- On Stationary-Point Hitting Time and Ergodicity of Stochastic Gradient Langevin Dynamics
- An Exponential Improvement on the Memorization Capacity of Deep Threshold Networks
- Provable Memorization via Deep Neural Networks using Sub-linear Parameters
- Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes
- Deepened Graph Auto-Encoders Help Stabilize and Enhance Link Prediction
- Spectrum concentration in deep residual learning: a free probability approach
- Optimal Function Approximation with Relu Neural Networks
- Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
- Noether: The More Things Change, the More Stay the Same
- Global Convergence of Gradient Descent for Deep Linear Residual Networks
- Understanding Global Loss Landscape of One-hidden-layer ReLU Networks, Part 1: Theory
- Ridge Regression with Over-Parametrized Two-Layer Networks Converge to Ridgelet Spectrum
- Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient Descent
- Identity Connections in Residual Nets Improve Noise Stability
- Fuzzy Logic Interpretation of Quadratic Networks
- ResNEsts and DenseNEsts: Block-based DNN Models with Improved Representation Guarantees
- Error estimate for a universal function approximator of ReLU network with a local connection
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks
- When Are Solutions Connected in Deep Networks?
- Deep Neural Networks Are Congestion Games: From Loss Landscape to Wardrop Equilibrium and Beyond
- Fine-grained Optimization of Deep Neural Networks
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations