Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
arXiv:1808.06866
Abstract
This paper proposed a Soft Filter Pruning (SFP) method to accelerate the inference procedure of deep Convolutional Neural Networks (CNNs). Specifically, the proposed SFP enables the pruned filters to be updated when training the model after pruning. SFP has two advantages over previous works: (1) Larger model capacity. Updating previously pruned filters provides our approach with larger optimization space than fixing the filters to zero. Therefore, the network trained by our method has a larger model capacity to learn from the training data. (2) Less dependence on the pre-trained model. Large capacity enables SFP to train from scratch and prune the model simultaneously. In contrast, previous filter pruning methods should be conducted on the basis of the pre-trained model to guarantee their performance. Empirically, SFP from scratch outperforms the previous filter pruning methods. Moreover, our approach has been demonstrated effective for many advanced CNN architectures. Notably, on ILSCRC-2012, SFP reduces more than 42% FLOPs on ResNet-101 with even 0.2% top-5 accuracy improvement, which has advanced the state-of-the-art. Code is publicly available on GitHub: https://github.com/he-y/soft-filter-pruning
Accepted to IJCAI 2018
References in corpus (9)
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Trained Ternary Quantization
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Learning Structured Sparsity in Deep Neural Networks
- Convolutional neural networks with low-rank regularization
- Channel Pruning for Accelerating Very Deep Neural Networks
- More is Less: A More Complicated Network with Less Inference Complexity
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
Cited by in corpus (47)
- DAIS: Automatic Channel Pruning via Differentiable Annealing Indicator Search
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions
- Hybrid Pruning: Thinner Sparse Networks for Fast Inference on Edge Devices
- Locally Free Weight Sharing for Network Width Search
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
- EagleEye: Fast Sub-net Evaluation for Efficient Neural Network Pruning
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- PAMS: Quantized Super-Resolution via Parameterized Max Scale
- Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
- Rethinking Class-Discrimination Based CNN Channel Pruning
- Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
- CUP: Cluster Pruning for Compressing Deep Neural Networks
- DSA: More Efficient Budgeted Pruning via Differentiable Sparsity Allocation
- Toward Compact Deep Neural Networks via Energy-Aware Pruning
- PENNI: Pruned Kernel Sharing for Efficient CNN Inference
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- Neural Network Compression Via Sparse Optimization
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification
- How Not to Give a FLOP: Combining Regularization and Pruning for Efficient Inference
- Dynamic Sparse Graph for Efficient Deep Learning
- A Generalized and Robust Method Towards Practical Gaze Estimation on Smart Phone
- Topological Insights into Sparse Neural Networks
- Dynamic Group Convolution for Accelerating Convolutional Neural Networks
- SASL: Saliency-Adaptive Sparsity Learning for Neural Network Acceleration
- Progressive Learning of Low-Precision Networks
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- Pruning Filter in Filter
- PFGDF: Pruning Filter via Gaussian Distribution Feature for Deep Neural Networks Acceleration
- A Feature-map Discriminant Perspective for Pruning Deep Neural Networks
- BCNet: Searching for Network Width with Bilaterally Coupled Network
- AdaPruner: Adaptive Channel Pruning and Effective Weights Inheritance
- BWCP: Probabilistic Learning-to-Prune Channels for ConvNets via Batch Whitening
- Filter Pruning using Hierarchical Group Sparse Regularization for Deep Convolutional Neural Networks
- ASCAI: Adaptive Sampling for acquiring Compact AI
- ViP: Virtual Pooling for Accelerating CNN-based Image Classification and Object Detection
- Width Transfer: On the (In)variance of Width Optimization
- Modulating Regularization Frequency for Efficient Compression-Aware Model Training
- Class-Discriminative CNN Compression
- Channel Planting for Deep Neural Networks using Knowledge Distillation
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
- Neural Plasticity Networks
- Out-of-the-box channel pruned networks