Learning Deep ResNet Blocks Sequentially using Boosting Theory
arXiv:1706.04964
Abstract
Deep neural networks are known to be difficult to train due to the instability of back-propagation. A deep \emph{residual network} (ResNet) with identity loops remedies this by stabilizing gradient computations. We prove a boosting theory for the ResNet architecture. We construct weak module classifiers, each contains two of the layers, such that the combined strong learner is a ResNet. Therefore, we introduce an alternative Deep ResNet training algorithm, \emph{BoostResNet}, which is particularly suitable in non-differentiable architectures. Our proposed algorithm merely requires a sequential training of "shallow ResNets" which are inexpensive. We prove that the training error decays exponentially with the depth if the \emph{weak module classifiers} that we train perform slightly better than some weak baseline. In other words, we propose a weak learning condition and prove a boosting theory for ResNet under the weak learning condition. Our results apply to general multi-class ResNets. A generalization error bound based on margin theory is proved and suggests ResNet's resistant to overfitting under network with norm bounded weights.
Accepted to ICML 2018
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- ADADELTA: An Adaptive Learning Rate Method
- Deep Residual Learning for Image Recognition
- Residual Networks Behave Like Ensembles of Relatively Shallow Networks
- Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
- AdaNet: Adaptive Structural Learning of Artificial Neural Networks
Cited by in corpus (12)
- Residual Connections Encourage Iterative Inference
- Learning From Noisy Labels By Regularized Estimation Of Annotator Confusion
- Auto-Meta: Automated Gradient Based Meta Learner Search
- Greedy Layerwise Learning Can Scale to ImageNet
- Tensorial Neural Networks: Generalization of Neural Networks and Application to Model Compression
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- Local Boosting for Weakly-Supervised Learning
- Functional Gradient Boosting based on Residual Network Perception
- Boosting Offline Reinforcement Learning with Residual Generative Modeling
- ChainGAN: A sequential approach to GANs
- Error Autocorrelation Objective Function for Improved System Modeling
- Neighbourhood Distillation: On the benefits of non end-to-end distillation