Tensor Decomposition for Compressing Recurrent Neural Network
arXiv:1802.10410
Abstract
In the machine learning fields, Recurrent Neural Network (RNN) has become a popular architecture for sequential data modeling. However, behind the impressive performance, RNNs require a large number of parameters for both training and inference. In this paper, we are trying to reduce the number of parameters and maintain the expressive power from RNN simultaneously. We utilize several tensor decompositions method including CANDECOMP/PARAFAC (CP), Tucker decomposition and Tensor Train (TT) to re-parameterize the Gated Recurrent Unit (GRU) RNN. We evaluate all tensor-based RNNs performance on sequence modeling tasks with a various number of parameters. Based on our experiment results, TT-GRU achieved the best results in a various number of parameters compared to other decomposition methods.
Accepted at IJCNN 2018. Source code URL: https://github.com/androstj/tensor_rnn
References in corpus (9)
- Distilling the Knowledge in a Neural Network
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- On the difficulty of training Recurrent Neural Networks
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- Deep Learning with Limited Numerical Precision
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Compressing Recurrent Neural Network with Tensor Train
- Sequence-Level Knowledge Distillation
- Learning Compact Recurrent Neural Networks with Block-Term Tensor Decomposition