Dynamic Model Pruning with Feedback
arXiv:2006.07253
Abstract
Deep neural networks often have millions of parameters. This can hinder their deployment to low-end devices, not only due to high memory requirements but also because of increased latency at inference. We propose a novel model compression method that generates a sparse trained model without additional overhead: by allowing (i) dynamic allocation of the sparsity pattern and (ii) incorporating feedback signal to reactivate prematurely pruned weights we obtain a performant sparse model in one single training pass (retraining is not needed, but can further improve the performance). We evaluate our method on CIFAR-10 and ImageNet, and show that the obtained sparse models can reach the state-of-the-art performance of dense models. Moreover, their performance surpasses that of models generated by all previously proposed pruning schemes.
appearing at ICLR 2020
References in corpus (4)
Cited by in corpus (6)
- Consistent Sparse Deep Learning: Theory and Computation
- Sparse Deep Learning: A New Framework Immune to Local Traps and Miscalibration
- Neural Pruning via Growing Regularization
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
- Distributed Sparse SGD with Majority Voting
- Architecture Aware Latency Constrained Sparse Neural Networks