Recurrent Memory Networks for Language Modeling
arXiv:1601.01272
Abstract
Recurrent Neural Networks (RNN) have obtained excellent result in many natural language processing (NLP) tasks. However, understanding and interpreting the source of this success remains a challenge. In this paper, we propose Recurrent Memory Network (RMN), a novel RNN architecture, that not only amplifies the power of RNN but also facilitates our understanding of its internal functioning and allows us to discover underlying patterns in data. We demonstrate the power of RMN on language modeling and sentence completion tasks. On language modeling, RMN outperforms Long Short-Term Memory (LSTM) network on three large German, Italian, and English dataset. Additionally we perform in-depth analysis of various linguistic dimensions that RMN captures. On Sentence Completion Challenge, for which it is essential to capture sentence coherence, our RMN obtains 69.2% accuracy, surpassing the previous state-of-the-art by a large margin.
8 pages, 6 figures. Accepted at NAACL 2016
References in corpus (4)
Cited by in corpus (15)
- Recent Advances in Recurrent Neural Networks
- Long Short-Term Memory-Networks for Machine Reading
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Learning Python Code Suggestion with a Sparse Pointer Network
- Attentive Memory Networks: Efficient Machine Reading for Conversational Search
- Neural Language Generation: Formulation, Methods, and Evaluation
- Self-Attentive Residual Decoder for Neural Machine Translation
- Learning to Remember Translation History with a Continuous Cache
- The Importance of Being Recurrent for Modeling Hierarchical Structure
- Unsupervised Neural Hidden Markov Models
- Reflective Decoding Network for Image Captioning
- Joint Embedding Learning of Educational Knowledge Graphs
- Attention-based Memory Selection Recurrent Network for Language Modeling
- Meta-Learning a Dynamical Language Model
- Improving Neural Language Models by Segmenting, Attending, and Predicting the Future