Channel Pruning for Accelerating Very Deep Neural Networks
arXiv:1707.06168
Abstract
In this paper, we introduce a new channel pruning method to accelerate very deep convolutional neural networks.Given a trained CNN model, we propose an iterative two-step algorithm to effectively prune each layer, by a LASSO regression based channel selection and least square reconstruction. We further generalize this algorithm to multi-layer and multi-branch cases. Our method reduces the accumulated error and enhance the compatibility with various architectures. Our pruned VGG-16 achieves the state-of-the-art results by 5x speed-up along with only 0.3% increase of error. More importantly, our method is able to accelerate modern networks like ResNet, Xception and suffers only 1.4%, 1.0% accuracy loss under 2x speed-up respectively, which is significant. Code has been made publicly available.
To be appear at ICCV 2017
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Compressing Deep Convolutional Networks using Vector Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Structured Sparsity in Deep Neural Networks
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Fast Convolutional Nets With fbfft: A GPU Performance Evaluation
- Compact Deep Convolutional Neural Networks With Coarse Pruning
Cited by in corpus (39)
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Slimmable Neural Networks
- Searching for Low-Bit Weights in Quantized Neural Networks
- PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training
- Channel Compression: Rethinking Information Redundancy among Channels in CNN Architecture
- Learning to Prune Filters in Convolutional Neural Networks
- APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- Hardware-Centric AutoML for Mixed-Precision Quantization
- Accelerating Sparse Deep Neural Networks
- An Experimental Study of the Impact of Pre-training on the Pruning of a Convolutional Neural Network
- Fine-Grained Neural Architecture Search
- Pruning via Iterative Ranking of Sensitivity Statistics
- Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks
- BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
- GeneCAI: Genetic Evolution for Acquiring Compact AI
- Structured Pruning for Efficient ConvNets via Incremental Regularization
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- CNNPruner: Pruning Convolutional Neural Networks with Visual Analytics
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- Enabling Retrain-free Deep Neural Network Pruning using Surrogate Lagrangian Relaxation
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- CAE-ADMM: Implicit Bitrate Optimization via ADMM-based Pruning in Compressive Autoencoders
- Real-Time Object Tracking via Meta-Learning: Efficient Model Adaptation and One-Shot Channel Pruning
- Cramnet: Layer-wise Deep Neural Network Compression with Knowledge Transfer from a Teacher Network
- Generalized Bayesian Posterior Expectation Distillation for Deep Neural Networks
- Channel-Level Variable Quantization Network for Deep Image Compression
- Improving Efficiency in Neural Network Accelerator Using Operands Hamming Distance optimization
- CondenseNet V2: Sparse Feature Reactivation for Deep Networks
- A Layer Decomposition-Recomposition Framework for Neuron Pruning towards Accurate Lightweight Networks
- AutoPruning for Deep Neural Network with Dynamic Channel Masking
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- A Main/Subsidiary Network Framework for Simplifying Binary Neural Network
- DAC: Data-free Automatic Acceleration of Convolutional Networks
- Auto-Split: A General Framework of Collaborative Edge-Cloud AI
- Feature Statistics Guided Efficient Filter Pruning
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks