Learning Pruned Structure and Weights Simultaneously from Scratch: an Attention based Approach
arXiv:2111.02399
Abstract
As a deep learning model typically contains millions of trainable weights, there has been a growing demand for a more efficient network structure with reduced storage space and improved run-time efficiency. Pruning is one of the most popular network compression techniques. In this paper, we propose a novel unstructured pruning pipeline, Attention-based Simultaneous sparse structure and Weight Learning (ASWL). Unlike traditional channel-wise or weight-wise attention mechanism, ASWL proposed an efficient algorithm to calculate the pruning ratio through layer-wise attention for each layer, and both weights for the dense network and the sparse network are tracked so that the pruned structure is simultaneously learned from randomly initialized weights. Our experiments on MNIST, Cifar10, and ImageNet show that ASWL achieves superior pruning results in terms of accuracy, pruning ratio and operating efficiency when compared with state-of-the-art network pruning methods.
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- Sparse Networks from Scratch: Faster Training without Losing Performance
- Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
- Dynamic Model Pruning with Feedback
- Proving the Lottery Ticket Hypothesis: Pruning is All You Need
- Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
- EDropout: Energy-Based Dropout and Pruning of Deep Neural Networks
- Learned Threshold Pruning
- Fine-Pruning: Joint Fine-Tuning and Compression of a Convolutional Network with Bayesian Optimization
- Pruning Filter in Filter