Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon
arXiv:1705.07565
Abstract
How to develop slim and accurate deep neural networks has become crucial for real- world applications, especially for those employed in embedded systems. Though previous work along this research line has shown some promising results, most existing methods either fail to significantly compress a well-trained deep network or require a heavy retraining process for the pruned deep network to re-boost its prediction performance. In this paper, we propose a new layer-wise pruning method for deep neural networks. In our proposed method, parameters of each individual layer are pruned independently based on second order derivatives of a layer-wise error function with respect to the corresponding parameters. We prove that the final prediction performance drop after pruning is bounded by a linear combination of the reconstructed errors caused at each layer. Therefore, there is a guarantee that one only needs to perform a light retraining process on the pruned network to resume its original prediction performance. We conduct extensive experiments on benchmark datasets to demonstrate the effectiveness of our pruning method compared with several state-of-the-art baseline methods.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- Compressing Deep Convolutional Networks using Vector Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- Convolutional neural networks with low-rank regularization
- Training Skinny Deep Neural Networks with Iterative Hard Thresholding Methods
- The Incredible Shrinking Neural Network: New Perspectives on Learning Representations Through The Lens of Pruning
Cited by in corpus (10)
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Neural network relief: a pruning algorithm based on neural activity
- F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
- Characterising Across-Stack Optimisations for Deep Convolutional Neural Networks
- You are caught stealing my winning lottery ticket! Making a lottery ticket claim its ownership
- Deep Neural Compression Via Concurrent Pruning and Self-Distillation
- Re-Weighted Learning for Sparsifying Deep Neural Networks
- FeTa: A DCA Pruning Algorithm with Generalization Error Guarantees
- Revisiting hard thresholding for DNN pruning