Depth-Gated LSTM
arXiv:1508.03790
Abstract
In this short note, we present an extension of long short-term memory (LSTM) neural networks to using a depth gate to connect memory cells of adjacent layers. Doing so introduces a linear dependence between lower and upper layer recurrent units. Importantly, the linear dependence is gated through a gating function, which we call depth gate. This gate is a function of the lower layer memory cell, the input to and the past memory cell of this layer. We conducted experiments and verified that this new architecture of LSTMs was able to improve machine translation and language modeling performances.
Content presented in 2015 Jelinek Summer Workshop on Speech and Language Technology on August 14th 2015
Cited by in corpus (8)
- Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions
- Recent Advances in Recurrent Neural Networks
- Long Short-Term Memory-Networks for Machine Reading
- Attention with Intention for a Neural Network Conversation Model
- The unreasonable effectiveness of the forget gate
- Shorten Spatial-spectral RNN with Parallel-GRU for Hyperspectral Image Classification
- ToyArchitecture: Unsupervised Learning of Interpretable Models of the World
- Recurrent Neural Networks with Flexible Gates using Kernel Activation Functions