SkipNet: Learning Dynamic Routing in Convolutional Networks
arXiv:1711.09485
Abstract
While deeper convolutional networks are needed to achieve maximum accuracy in visual perception tasks, for many inputs shallower networks are sufficient. We exploit this observation by learning to skip convolutional layers on a per-input basis. We introduce SkipNet, a modified residual network, that uses a gating network to selectively skip convolutional blocks based on the activations of the previous layer. We formulate the dynamic skipping problem in the context of sequential decision making and propose a hybrid learning algorithm that combines supervised learning and reinforcement learning to address the challenges of non-differentiable skipping decisions. We show SkipNet reduces computation by 30-90% while preserving the accuracy of the original model on four benchmark datasets and outperforms the state-of-the-art dynamic networks and static compression methods. We also qualitatively evaluate the gating policy to reveal a relationship between image scale and saliency and the number of layers skipped.
ECCV 2018 Camera ready version. Code is available at https://github.com/ucbdrive/skipnet
References in corpus (1)
Cited by in corpus (14)
- Slimmable Neural Networks
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- A Survey of FPGA-Based Neural Network Accelerator
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Competitive Inner-Imaging Squeeze and Excitation for Residual Network
- Improved Techniques for Training Adaptive Deep Networks
- Anytime Inference with Distilled Hierarchical Neural Ensembles
- Learning Anytime Predictions in Neural Networks via Adaptive Loss Balancing
- LAP-Net: Adaptive Features Sampling via Learning Action Progression for Online Action Detection
- TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning
- SGAD: Soft-Guided Adaptively-Dropped Neural Network
- Graceful Degradation and Related Fields
- Dynamic Multi-path Neural Network
- Efficient Video Understanding via Layered Multi Frame-Rate Analysis