The Shattered Gradients Problem: If resnets are the answer, then what is the question?
arXiv:1702.08591
Abstract
A long-standing obstacle to progress in deep learning is the problem of vanishing and exploding gradients. Although, the problem has largely been overcome via carefully constructed initializations and batch normalization, architectures incorporating skip-connections such as highway and resnets perform much better than standard feedforward architectures despite well-chosen initialization and batch normalization. In this paper, we identify the shattered gradients problem. Specifically, we show that the correlation between gradients in standard feedforward networks decays exponentially with depth resulting in gradients that resemble white noise whereas, in contrast, the gradients in architectures with skip-connections are far more resistant to shattering, decaying sublinearly. Detailed empirical evidence is presented in support of the analysis, on both fully-connected networks and convnets. Finally, we present a new "looks linear" (LL) initialization that prevents shattering, with preliminary experiments showing the new initialization allows to train very deep networks without the addition of skip-connections.
ICML 2017, final version
References in corpus (8)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Residual Networks Behave Like Ensembles of Relatively Shallow Networks
- Highway and Residual Networks learn Unrolled Iterative Estimation
- DiracNets: Training Very Deep Neural Networks Without Skip-Connections
- Random Walk Initialization for Training Very Deep Feedforward Networks
- Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks
- Deep Online Convex Optimization with Gated Games
Cited by in corpus (38)
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Deep Learning for Single Image Super-Resolution: A Brief Review
- High-Performance Large-Scale Image Recognition Without Normalization
- How Does Batch Normalization Help Optimization?
- Optimization for deep learning: theory and algorithms
- Investigating the fidelity of explainable artificial intelligence methods for applications of convolutional neural networks in geoscience
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- A Mean Field Theory of Batch Normalization
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- Understanding Neural Code Intelligence Through Program Simplification
- White Box Methods for Explanations of Convolutional Neural Networks in Image Classification Tasks
- NormFormer: Improved Transformer Pretraining with Extra Normalization
- Empirical Studies on the Properties of Linear Regions in Deep Neural Networks
- Continuous-in-Depth Neural Networks
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
- Hybrid Batch Attacks: Finding Black-box Adversarial Examples with Limited Queries
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Positional Normalization
- Regularizing activations in neural networks via distribution matching with the Wasserstein metric
- The Nonlinearity Coefficient - Predicting Generalization in Deep Neural Networks
- Initialization of ReLUs for Dynamical Isometry
- ResNet After All? Neural ODEs and Their Numerical Solution
- The Nonlinearity Coefficient - A Practical Guide to Neural Architecture Design
- A Deep Conditioning Treatment of Neural Networks
- Learning Residue-Aware Correlation Filters and Refining Scale Estimates with the GrabCut for Real-Time UAV Tracking
- Attention-Based Clustering: Learning a Kernel from Context
- Extreme Memorization via Scale of Initialization
- The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization
- On the Bias-Variance Tradeoff: Textbooks Need an Update
- Augmented Shortcuts for Vision Transformers
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- Large-Scale Meta-Learning with Continual Trajectory Shifting
- Separating the Effects of Batch Normalization on CNN Training Speed and Stability Using Classical Adaptive Filter Theory
- Conflicting Bundles: Adapting Architectures Towards the Improved Training of Deep Neural Networks
- Adaptive Weight Decay for Deep Neural Networks
- Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
- On the Existence of Universal Lottery Tickets