Learning Intrinsic Sparse Structures within Long Short-Term Memory
arXiv:1709.05027
Abstract
Model compression is significant for the wide adoption of Recurrent Neural Networks (RNNs) in both user devices possessing limited resources and business clusters requiring quick responses to large-scale service requests. This work aims to learn structurally-sparse Long Short-Term Memory (LSTM) by reducing the sizes of basic structures within LSTM units, including input updates, gates, hidden states, cell states and outputs. Independently reducing the sizes of basic structures can result in inconsistent dimensions among them, and consequently, end up with invalid LSTM units. To overcome the problem, we propose Intrinsic Sparse Structures (ISS) in LSTMs. Removing a component of ISS will simultaneously decrease the sizes of all basic structures by one and thereby always maintain the dimension consistency. By learning ISS within LSTM units, the obtained LSTMs remain regular while having much smaller basic structures. Based on group Lasso regularization, our method achieves 10.59x speedup without losing any perplexity of a language modeling of Penn TreeBank dataset. It is also successfully evaluated through a compact model with only 2.69M weights for machine Question Answering of SQuAD dataset. Our approach is successfully extended to non- LSTM RNNs, like Recurrent Highway Networks (RHNs). Our source code is publicly available at https://github.com/wenwei202/iss-rnns
Published in ICLR 2018 ( the Sixth International Conference on Learning Representations)
References in corpus (8)
- Distilling the Knowledge in a Neural Network
- Neural Architecture Search with Reinforcement Learning
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Recurrent Neural Network Regularization
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Structured Sparsity in Deep Neural Networks
- Quasi-Recurrent Neural Networks
- Nonparametric Neural Networks
Cited by in corpus (39)
- TinyLSTMs: Efficient Neural Speech Enhancement for Hearing Aids
- Block-Sparse Recurrent Neural Networks
- Balanced Sparsity for Efficient DNN Inference on GPU
- PruneTrain: Fast Neural Network Training by Dynamic Sparse Model Reconfiguration
- Crossbar-aware neural network pruning
- Spartus: A 9.4 TOp/s FPGA-based LSTM Accelerator Exploiting Spatio-Temporal Sparsity
- AutoPruner: An End-to-End Trainable Filter Pruning Method for Efficient Deep Model Inference
- On the Effectiveness of Low-Rank Matrix Factorization for LSTM Model Compression
- Inference skipping for more efficient real-time speech enhancement with parallel RNNs
- End-to-End Supermask Pruning: Learning to Prune Image Captioning Models
- Selfish Sparse RNN Training
- One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
- Growing Efficient Deep Networks by Structured Continuous Sparsification
- SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference
- Accelerating CNN Training by Pruning Activation Gradients
- Hardware-Guided Symbiotic Training for Compact, Accurate, yet Execution-Efficient LSTM
- Intrinsically Sparse Long Short-Term Memory Networks
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- AutoGrow: Automatic Layer Growing in Deep Convolutional Networks
- Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification
- Compressing Gradient Optimizers via Count-Sketches
- Self-Supervised GAN Compression
- DiabDeep: Pervasive Diabetes Diagnosis based on Wearable Medical Sensors and Efficient Neural Networks
- Compressing LSTM Networks by Matrix Product Operators
- Block-wise Dynamic Sparseness
- Efficient Sparse-Dense Matrix-Matrix Multiplication on GPUs Using the Customized Sparse Storage Format
- Streaming Voice Query Recognition using Causal Convolutional Recurrent Neural Networks
- AntMan: Sparse Low-Rank Compression to Accelerate RNN inference
- Structural sparsification for Far-field Speaker Recognition with GNA
- DARB: A Density-Aware Regular-Block Pruning for Deep Neural Networks
- Bayesian Sparsification of Gated Recurrent Neural Networks
- Spectral Pruning for Recurrent Neural Networks
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training
- Block-term Tensor Neural Networks
- Learning to Actively Reduce Memory Requirements for Robot Control Tasks
- Neural network compression via learnable wavelet transforms
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- Approximate Random Dropout