Larger-Context Language Modelling
arXiv:1511.03729
Abstract
In this work, we propose a novel method to incorporate corpus-level discourse information into language modelling. We call this larger-context language model. We introduce a late fusion approach to a recurrent language model based on long short-term memory units (LSTM), which helps the LSTM unit keep intra-sentence dependencies and inter-sentence dependencies separate from each other. Through the evaluation on three corpora (IMDB, BBC, and PennTree Bank), we demon- strate that the proposed model improves perplexity significantly. In the experi- ments, we evaluate the proposed approach while varying the number of context sentences and observe that the proposed late fusion is superior to the usual way of incorporating additional inputs to the LSTM. By analyzing the trained larger- context language model, we discover that content words, including nouns, adjec- tives and verbs, benefit most from an increasing number of context sentences. This analysis suggests that larger-context language model improves the unconditional language model by capturing the theme of a document better and more easily.
References in corpus (5)
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- The Goldilocks Principle: Reading Children's Books with Explicit Memory Representations
- Learning Longer Memory in Recurrent Neural Networks
Cited by in corpus (20)
- Exploring the Limits of Language Modeling
- Context-aware Natural Language Generation with Recurrent Neural Networks
- Document Context Language Models
- Natural Language Understanding with Distributed Representation
- A Survey on Neural Network Language Models
- Deep Multi-Task Learning with Shared Memory
- Capitalization and Punctuation Restoration: a Survey
- Multi-Modal Deep Learning for Credit Rating Prediction Using Text and Numerical Data Streams
- Unbounded cache model for online language modeling with open vocabulary
- Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
- Reference-Aware Language Models
- PopMAG: Pop Music Accompaniment Generation
- Contextual Text Style Transfer
- Cross-Utterance Language Models with Acoustic Error Sampling
- Cross-Attention End-to-End ASR for Two-Party Conversations
- Dialog-context aware end-to-end speech recognition
- Learning Simpler Language Models with the Differential State Framework
- Better Long-Range Dependency By Bootstrapping A Mutual Information Regularizer
- Acoustic-to-Word Models with Conversational Context Information
- Dialog Context Language Modeling with Recurrent Neural Networks