Exploring Sparsity in Recurrent Neural Networks
arXiv:1704.05119
Abstract
Recurrent Neural Networks (RNN) are widely used to solve a variety of problems and as the quantity of data and the amount of available compute have increased, so have model sizes. The number of parameters in recent state-of-the-art networks makes them hard to deploy, especially on mobile phones and embedded devices. The challenge is due to both the size of the model and the time it takes to evaluate it. In order to deploy these RNNs efficiently, we propose a technique to reduce the parameters of a network by pruning weights during the initial training of the network. At the end of training, the parameters of the network are sparse while accuracy is still close to the original dense neural network. The network size is reduced by 8x and the time required to train the model remains constant. Additionally, we can prune a larger dense network to achieve better than baseline performance while still reducing the total number of parameters significantly. Pruning RNNs reduces the size of the model and can also help achieve significant inference time speed-up using sparse matrix multiply. Benchmarks show that using our technique model size can be reduced by 90% and speed-up is around 2x to 7x.
Published as a conference paper at ICLR 2017
References in corpus (5)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Speeding up Convolutional Neural Networks with Low Rank Expansions
Cited by in corpus (73)
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- The State of Sparsity in Deep Neural Networks
- Sparse Networks from Scratch: Faster Training without Losing Performance
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- Recurrent Neural Networks: An Embedded Computing Perspective
- Learning Intrinsic Sparse Structures within Long Short-Term Memory
- Rigging the Lottery: Making All Tickets Winners
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- Soft Threshold Weight Reparameterization for Learnable Sparsity
- FastGRNN: A Fast, Accurate, Stable and Tiny Kilobyte Sized Gated Recurrent Neural Network
- Block-Sparse Recurrent Neural Networks
- What Do Compressed Deep Neural Networks Forget?
- Efficient Neural Audio Synthesis
- Neural Network Distiller: A Python Package For DNN Compression Research
- Characterising Bias in Compressed Models
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Looking GLAMORous: Vehicle Re-Id in Heterogeneous Cameras Networks with Global and Local Attention
- DeepTwist: Learning Model Compression via Occasional Weight Distortion
- Grounded Recurrent Neural Networks
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Grow and Prune Compact, Fast, and Accurate LSTMs
- Hierarchical Block Sparse Neural Networks
- One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation
- Accelerating Sparse Deep Neural Networks
- Bayesian Sparsification of Recurrent Neural Networks
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
- Retraining-Based Iterative Weight Quantization for Deep Neural Networks
- Ternary Hybrid Neural-Tree Networks for Highly Constrained IoT Applications
- Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator
- Growing Efficient Deep Networks by Structured Continuous Sparsification
- Hardware-Guided Symbiotic Training for Compact, Accurate, yet Execution-Efficient LSTM
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Transformed Regularization for Learning Sparse Deep Neural Networks
- Intrinsically Sparse Long Short-Term Memory Networks
- On improving deep learning generalization with adaptive sparse connectivity
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- Topological Insights into Sparse Neural Networks
- GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking
- Enhancing the Regularization Effect of Weight Pruning in Artificial Neural Networks
- AntMan: Sparse Low-Rank Compression to Accelerate RNN inference
- Efficient Sparse-Dense Matrix-Matrix Multiplication on GPUs Using the Customized Sparse Storage Format
- Block-wise Dynamic Sparseness
- Image Captioning with Sparse Recurrent Neural Network
- DiabDeep: Pervasive Diabetes Diagnosis based on Wearable Medical Sensors and Efficient Neural Networks
- Dynamical Phases and Resonance Phenomena in Information-Processing Recurrent Neural Networks
- Small, Accurate, and Fast Vehicle Re-ID on the Edge: the SAFR Approach
- Are wider nets better given the same number of parameters?
- Structural sparsification for Far-field Speaker Recognition with GNA
- Reinforcement Learning with Chromatic Networks for Compact Architecture Search
- Bayesian Sparsification of Gated Recurrent Neural Networks
- Machine Learning Approach for Transforming Scattering Parameters to Complex Permittivity
- FeatherTTS: Robust and Efficient attention based Neural TTS
- Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation
- Magnitude and Uncertainty Pruning Criterion for Neural Networks
- HALO: Learning to Prune Neural Networks with Shrinkage
- Block-term Tensor Neural Networks
- Neural networks adapting to datasets: learning network size and topology
- Neural network compression via learnable wavelet transforms
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- Weight, Block or Unit? Exploring Sparsity Tradeoffs for Speech Enhancement on Tiny Neural Accelerators
- Sparse Multi-Family Deep Scattering Network
- Sparsity Emerges Naturally in Neural Language Models
- A Survey on Green Deep Learning
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training
- Training Sparse Neural Network by Constraining Synaptic Weight on Unit Lp Sphere
- SlimNets: An Exploration of Deep Model Compression and Acceleration