Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
arXiv:1709.06994
Abstract
In this paper, we propose a novel progressive parameter pruning method for Convolutional Neural Network acceleration, named Structured Probabilistic Pruning (SPP), which effectively prunes weights of convolutional layers in a probabilistic manner. Unlike existing deterministic pruning approaches, where unimportant weights are permanently eliminated, SPP introduces a pruning probability for each weight, and pruning is guided by sampling from the pruning probabilities. A mechanism is designed to increase and decrease pruning probabilities based on importance criteria in the training process. Experiments show that, with 4x speedup, SPP can accelerate AlexNet with only 0.3% loss of top-5 accuracy and VGG-16 with 0.8% loss of top-5 accuracy in ImageNet classification. Moreover, SPP can be directly applied to accelerate multi-branch CNN networks, such as ResNet, without specific adaptations. Our 2x speedup ResNet-50 only suffers 0.8% loss of top-5 accuracy on ImageNet. We further show the effectiveness of SPP on transfer learning tasks.
CNN model acceleration, 13 pages, 6 figures, accepted by Proceedings of the British Machine Vision Conference (BMVC), 2018 oral
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Improving neural networks by preventing co-adaptation of feature detectors
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- cuDNN: Efficient Primitives for Deep Learning
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Structured Sparsity in Deep Neural Networks
- Channel Pruning for Accelerating Very Deep Neural Networks
- Bayesian Compression for Deep Learning
- Fast Convolutional Nets With fbfft: A GPU Performance Evaluation
- Compact Deep Convolutional Neural Networks With Coarse Pruning
Cited by in corpus (18)
- Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
- Discrimination-aware Network Pruning for Deep Model Compression
- Balanced Sparsity for Efficient DNN Inference on GPU
- Learning Sparse Networks Using Targeted Dropout
- Layer-compensated Pruning for Resource-constrained Convolutional Neural Networks
- Hybrid Pruning: Thinner Sparse Networks for Fast Inference on Edge Devices
- GeneCAI: Genetic Evolution for Acquiring Compact AI
- Taxonomy of Saliency Metrics for Channel Pruning
- Structured Pruning for Efficient ConvNets via Incremental Regularization
- Collaborative Distillation for Ultra-Resolution Universal Style Transfer
- Pruning Deep Neural Networks using Partial Least Squares
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- Network Automatic Pruning: Start NAP and Take a Nap
- Three Dimensional Convolutional Neural Network Pruning with Regularization-Based Method
- ASCAI: Adaptive Sampling for acquiring Compact AI
- Channel-wise pruning of neural networks with tapering resource constraint
- Communication-Computation Efficient Device-Edge Co-Inference via AutoML
- Directed-Weighting Group Lasso for Eltwise Blocked CNN Pruning