Soft Weight-Sharing for Neural Network Compression
arXiv:1702.04008
Abstract
The success of deep learning in numerous application domains created the de- sire to run and train them on mobile devices. This however, conflicts with their computationally, memory and energy intense nature, leading to a growing interest in compression. Recent work by Han et al. (2015a) propose a pipeline that involves retraining, pruning and quantization of neural network weights, obtaining state-of-the-art compression rates. In this paper, we show that competitive compression rates can be achieved by using a version of soft weight-sharing (Nowlan & Hinton, 1992). Our method achieves both quantization and pruning in one simple (re-)training procedure. This point of view also exposes the relation between compression and the minimum description length (MDL) principle.
ICLR2017
Cited by in corpus (31)
- Bayesian Compression for Deep Learning
- Practical Lossless Compression with Latent Variables using Bits Back Coding
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- Pruning Convolutional Neural Networks with Self-Supervision
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
- Learned Threshold Pruning
- Resource-Efficient Neural Networks for Embedded Systems
- SeReNe: Sensitivity based Regularization of Neurons for Structured Sparsity in Neural Networks
- LSQ+: Improving low-bit quantization through learnable offsets and better initialization
- Bayesian Tensorized Neural Networks with Automatic Rank Selection
- Iteratively Training Look-Up Tables for Network Quantization
- Uncertainty Quantification for Sparse Deep Learning
- A Highly Parallel FPGA Implementation of Sparse Neural Network Training
- Improved Bayesian Compression
- HiLLoC: Lossless Image Compression with Hierarchical Latent Variable Models
- Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
- DKM: Differentiable K-Means Clustering Layer for Neural Network Compression
- Convergence of a Relaxed Variable Splitting Method for Learning Sparse Neural Networks via , and transformed- Penalties
- Train-by-Reconnect: Decoupling Locations of Weights from their Values
- FALCON: Lightweight and Accurate Convolution
- Exploring Weight Symmetry in Deep Neural Networks
- ESPN: Extremely Sparse Pruned Networks
- DP-Net: Dynamic Programming Guided Deep Neural Network Compression
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- Lossless Compression with Latent Variable Models
- Multi-Glimpse Network: A Robust and Efficient Classification Architecture based on Recurrent Downsampled Attention
- Distilling Image Classifiers in Object Detectors
- Resource-Efficient Speech Mask Estimation for Multi-Channel Speech Enhancement
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent Method for Learning Sparse Weight Binarized Activation Neural Networks
- Cascade Weight Shedding in Deep Neural Networks: Benefits and Pitfalls for Network Pruning