Residual Connections Encourage Iterative Inference
arXiv:1710.04773
Abstract
Residual networks (Resnets) have become a prominent architecture in deep learning. However, a comprehensive understanding of Resnets is still a topic of ongoing research. A recent view argues that Resnets perform iterative refinement of features. We attempt to further expose properties of this aspect. To this end, we study Resnets both analytically and empirically. We formalize the notion of iterative refinement in Resnets by showing that residual connections naturally encourage features of residual blocks to move along the negative gradient of loss as we go from one block to the next. In addition, our empirical analysis suggests that Resnets are able to perform both representation learning and iterative refinement. In general, a Resnet block tends to concentrate representation learning behavior in the first few layers while higher layers perform iterative refinement of features. Finally we observe that sharing residual layers naively leads to representation explosion and counterintuitively, overfitting, and we show that simple existing strategies can help alleviating this problem.
First two authors contributed equally. Published in ICLR 2018
References in corpus (2)
Cited by in corpus (12)
- A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
- Going in circles is the way forward: the role of recurrence in visual inference
- Multi-level Residual Networks from Dynamical Systems View
- NAIS-Net: Stable Deep Networks from Non-Autonomous Differential Equations
- CASI: A Convolutional Neural Network Approach for Shell Identification
- IamNN: Iterative and Adaptive Mobile Neural Network for Efficient Image Classification
- Learning Implicitly Recurrent CNNs Through Parameter Sharing
- SGAD: Soft-Guided Adaptively-Dropped Neural Network
- Hidden-Fold Networks: Random Recurrent Residuals Using Sparse Supermasks
- Functional Gradient Boosting based on Residual Network Perception
- Neural Function Modules with Sparse Arguments: A Dynamic Approach to Integrating Information across Layers
- An investigation of model-free planning