Skip Connections Eliminate Singularities
arXiv:1701.09175
Abstract
Skip connections made the training of very deep networks possible and have become an indispensable component in a variety of neural architectures. A completely satisfactory explanation for their success remains elusive. Here, we present a novel explanation for the benefits of skip connections in training very deep networks. The difficulty of training deep networks is partly due to the singularities caused by the non-identifiability of the model. Several such singularities have been identified in previous works: (i) overlap singularities caused by the permutation symmetry of nodes in a given layer, (ii) elimination singularities corresponding to the elimination, i.e. consistent deactivation, of nodes, (iii) singularities generated by the linear dependence of the nodes. These singularities cause degenerate manifolds in the loss landscape that slow down learning. We argue that skip connections eliminate these singularities by breaking the permutation symmetry of nodes, by reducing the possibility of node elimination and by making the nodes less linearly dependent. Moreover, for typical initializations, skip connections move the network away from the "ghosts" of these singularities and sculpt the landscape around them to alleviate the learning slow-down. These hypotheses are supported by evidence from simplified models, as well as from experiments with deep networks trained on real-world datasets.
Published as a conference paper at ICLR 2018
Cited by in corpus (25)
- Automatically designing CNN architectures using genetic algorithm for image classification
- Micro-Batch Training with Batch-Channel Normalization and Weight Standardization
- Deep SCNN-based Real-time Object Detection for Self-driving Vehicles Using LiDAR Temporal Data
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted Surgery
- Approximation spaces of deep neural networks
- Theory-Inspired Path-Regularized Differential Network Architecture Search
- Automatically Evolving CNN Architectures Based on Blocks
- Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
- Rethinking Normalization and Elimination Singularity in Neural Networks
- OMPQ: Orthogonal Mixed Precision Quantization
- ST-MTL: Spatio-Temporal Multitask Learning Model to Predict Scanpath While Tracking Instruments in Robotic Surgery
- On the Principle of Least Symmetry Breaking in Shallow ReLU Models
- Information-Theoretic Local Minima Characterization and Regularization
- FADNet++: Real-Time and Accurate Disparity Estimation with Configurable Networks
- LSALSA: Accelerated Source Separation via Learned Sparse Coding
- Learning Residue-Aware Correlation Filters and Refining Scale Estimates with the GrabCut for Real-Time UAV Tracking
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- Identity Connections in Residual Nets Improve Noise Stability
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- Neural Network Training Techniques Regularize Optimization Trajectory: An Empirical Study
- External-Memory Networks for Low-Shot Learning of Targets in Forward-Looking-Sonar Imagery
- Multi-Model Learning for Real-Time Automotive Semantic Foggy Scene Understanding via Domain Adaptation
- W-Cell-Net: Multi-frame Interpolation of Cellular Microscopy Videos
- Veritatem Dies Aperit- Temporally Consistent Depth Prediction Enabled by a Multi-Task Geometric and Semantic Scene Understanding Approach