Factorization tricks for LSTM networks
arXiv:1703.10722
Abstract
We present two simple ways of reducing the number of parameters and accelerating the training of large Long Short-Term Memory (LSTM) networks: the first one is "matrix factorization by design" of LSTM matrix into the product of two smaller matrices, and the second one is partitioning of LSTM matrix, its inputs and states into the independent groups. Both approaches allow us to train large LSTM networks significantly faster to the near state-of the art perplexity while using significantly less RNN parameters.
accepted to ICLR 2017 Workshop
References in corpus (3)
Cited by in corpus (22)
- Depthwise Separable Convolutions for Neural Machine Translation
- Exploring Interpretable LSTM Neural Networks over Multi-Variable Data
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Federated Learning Meets Natural Language Processing: A Survey
- Compressing RNNs for IoT devices by 15-38x using Kronecker Products
- On the Effectiveness of Low-Rank Matrix Factorization for LSTM Model Compression
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- Multi-variable LSTM neural network for autoregressive exogenous model
- Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs
- Faster Neural Network Training with Approximate Tensor Operations
- Intrinsically Sparse Long Short-Term Memory Networks
- Driving with Data: Modeling and Forecasting Vehicle Fleet Maintenance in Detroit
- Trace norm regularization and faster inference for embedded speech recognition RNNs
- Compressing Language Models using Doped Kronecker Products
- Low-Rank RNN Adaptation for Context-Aware Language Modeling
- Driving with Data in the Motor City: Mining and Modeling Vehicle Fleet Maintenance Data
- Low Rank Factorization for Compact Multi-Head Self-Attention
- A Lightweight Recurrent Network for Sequence Modeling
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices
- In-training Matrix Factorization for Parameter-frugal Neural Machine Translation
- Adversarially Robust and Explainable Model Compression with On-Device Personalization for Text Classification