Channel Pruning for Accelerating Very Deep Neural Networks
arXiv:1707.06168
Abstract
In this paper, we introduce a new channel pruning method to accelerate very deep convolutional neural networks.Given a trained CNN model, we propose an iterative two-step algorithm to effectively prune each layer, by a LASSO regression based channel selection and least square reconstruction. We further generalize this algorithm to multi-layer and multi-branch cases. Our method reduces the accumulated error and enhance the compatibility with various architectures. Our pruned VGG-16 achieves the state-of-the-art results by 5x speed-up along with only 0.3% increase of error. More importantly, our method is able to accelerate modern networks like ResNet, Xception and suffers only 1.4%, 1.0% accuracy loss under 2x speed-up respectively, which is significant. Code has been made publicly available.
To be appear at ICCV 2017
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Compressing Deep Convolutional Networks using Vector Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Structured Sparsity in Deep Neural Networks
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Fast Convolutional Nets With fbfft: A GPU Performance Evaluation
- Compact Deep Convolutional Neural Networks With Coarse Pruning
Cited by in corpus (6)
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training
- Learning to Prune Filters in Convolutional Neural Networks
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Cramnet: Layer-wise Deep Neural Network Compression with Knowledge Transfer from a Teacher Network
- CAE-ADMM: Implicit Bitrate Optimization via ADMM-based Pruning in Compressive Autoencoders