Highway and Residual Networks learn Unrolled Iterative Estimation
arXiv:1612.07771
Abstract
The past year saw the introduction of new architectures such as Highway networks and Residual networks which, for the first time, enabled the training of feedforward networks with dozens to hundreds of layers using simple gradient descent. While depth of representation has been posited as a primary reason for their success, there are indications that these architectures defy a popular view of deep learning as a hierarchical computation of increasingly abstract features at each layer. In this report, we argue that this view is incomplete and does not adequately explain several recent findings. We propose an alternative viewpoint based on unrolled iterative estimation -- a group of successive layers iteratively refine their estimates of the same features instead of computing an entirely new representation. We demonstrate that this viewpoint directly leads to the construction of Highway and Residual networks. Finally we provide preliminary experiments to discuss the similarities and differences between the two architectures.
10 + 4 pages, accepted for ICLR 2017
Cited by in corpus (15)
- Rethinking Architecture Selection in Differentiable NAS
- Learning Implicitly Recurrent CNNs Through Parameter Sharing
- Recurrence along Depth: Deep Convolutional Neural Networks with Recurrent Layer Aggregation
- ODE Transformer: An Ordinary Differential Equation-Inspired Model for Neural Machine Translation
- Graph Classification via Deep Learning with Virtual Nodes
- The Nonlinearity Coefficient - A Practical Guide to Neural Architecture Design
- Layer Folding: Neural Network Depth Reduction using Activation Linearization
- Exploring Weight Symmetry in Deep Neural Networks
- Identity Connections in Residual Nets Improve Noise Stability
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- Learning Light-Weight Translation Models from Deep Transformer
- Uncertainty Quantification in Deep Residual Neural Networks
- E2-Train: Training State-of-the-art CNNs with Over 80% Energy Savings
- Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining
- OrthoReg: Robust Network Pruning Using Orthonormality Regularization