More is Less: A More Complicated Network with Less Inference Complexity
arXiv:1703.08651
Abstract
In this paper, we present a novel and general network structure towards accelerating the inference process of convolutional neural networks, which is more complicated in network structure yet with less inference complexity. The core idea is to equip each original convolutional layer with another low-cost collaborative layer (LCCL), and the element-wise multiplication of the ReLU outputs of these two parallel layers produces the layer-wise output. The combined layer is potentially more discriminative than the original convolutional layer, and its inference is faster for two reasons: 1) the zero cells of the LCCL feature maps will remain zero after element-wise multiplication, and thus it is safe to skip the calculation of the corresponding high-cost convolution in the original convolutional layer, 2) LCCL is very fast if it is implemented as a 1*1 convolution or only a single filter shared by all channels. Extensive experiments on the CIFAR-10, CIFAR-100 and ILSCRC-2012 benchmarks show that our proposed network structure can accelerate the inference process by 32\% on average with negligible performance drop.
This paper has been accepted by the IEEE CVPR 2017
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Deep Learning with Limited Numerical Precision
- Pruning Filters for Efficient ConvNets
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Spatially-sparse convolutional neural networks
Cited by in corpus (26)
- Bayesian Compression for Deep Learning
- Dynamic Channel Pruning: Feature Boosting and Suppression
- Compressing Neural Networks using the Variational Information Bottleneck
- Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
- Style Aggregated Network for Facial Landmark Detection
- SkipNet: Learning Dynamic Routing in Convolutional Networks
- DAIS: Automatic Channel Pruning via Differentiable Annealing Indicator Search
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Dynamic Neural Networks: A Survey
- Deep Reinforcement Learning with Population-Coded Spiking Neural Network for Continuous Control
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- Learning to Prune Filters in Convolutional Neural Networks
- Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling
- Locally Free Weight Sharing for Network Width Search
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- SBNet: Sparse Blocks Network for Fast Inference
- Fine-Grained Neural Architecture Search
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- Computation on Sparse Neural Networks: an Inspiration for Future Hardware
- ENAS4D: Efficient Multi-stage CNN Architecture Search for Dynamic Inference
- Cross-Channel Intragroup Sparsity Neural Network
- Recurrent Residual Module for Fast Inference in Videos
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
- Softer Pruning, Incremental Regularization
- Online Filter Clustering and Pruning for Efficient Convnets