Learning Compact Recurrent Neural Networks
arXiv:1604.02594
Abstract
Recurrent neural networks (RNNs), including long short-term memory (LSTM) RNNs, have produced state-of-the-art results on a variety of speech recognition tasks. However, these models are often too large in size for deployment on mobile devices with memory and latency constraints. In this work, we study mechanisms for learning compact RNNs and LSTMs via low-rank factorizations and parameter sharing schemes. Our goal is to investigate redundancies in recurrent architectures where compression can be admitted without losing performance. A hybrid strategy of using structured matrices in the bottom layers and shared low-rank factors on the top layers is found to be particularly effective, reducing the parameters of a standard LSTM by 75%, at a small cost of 0.3% increase in WER, on a 2,000-hr English Voice Search task.
References in corpus (2)
Cited by in corpus (13)
- Learning Intrinsic Sparse Structures within Long Short-Term Memory
- Sequence-Level Knowledge Distillation
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- Learning Compact Recurrent Neural Networks with Block-Term Tensor Decomposition
- Grow and Prune Compact, Fast, and Accurate LSTMs
- Faster Neural Network Training with Approximate Tensor Operations
- Multiscale Hierarchical Convolutional Networks
- Intrinsically Sparse Long Short-Term Memory Networks
- GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- Training Recurrent Neural Networks as a Constraint Satisfaction Problem
- An Embedded Deep Learning based Word Prediction