Contextual LSTM (CLSTM) models for Large scale NLP tasks
arXiv:1602.06291
Abstract
Documents exhibit sequential structure at multiple levels of abstraction (e.g., sentences, paragraphs, sections). These abstractions constitute a natural hierarchy for representing the context in which to infer the meaning of words and larger fragments of text. In this paper, we present CLSTM (Contextual LSTM), an extension of the recurrent neural network LSTM (Long-Short Term Memory) model, where we incorporate contextual features (e.g., topics) into the model. We evaluate CLSTM on three specific NLP tasks: word prediction, next sentence selection, and sentence topic prediction. Results from experiments run on two corpora, English documents in Wikipedia and a subset of articles from a recent snapshot of English Google News, indicate that using both words and topics as features improves performance of the CLSTM models over baseline LSTM models for these tasks. For example on the next sentence selection task, we get relative accuracy improvements of 21% for the Wikipedia dataset and 18% for the Google News dataset. This clearly demonstrates the significant benefit of using context appropriately in natural language (NL) tasks. This has implications for a wide variety of NL applications like question answering, sentence completion, paraphrase generation, and next utterance prediction in dialog systems.
References in corpus (9)
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Text Understanding from Scratch
- Grammar as a Foreign Language
- Catching the Drift: Probabilistic Content Models, with Applications to Generation and Summarization
- Document Embedding with Paragraph Vectors
- Long Short-Term Memory Over Tree Structures
- Document Context Language Models
Cited by in corpus (38)
- Learning from History and Present: Next-item Recommendation via Discriminatively Exploiting User Behaviors
- Patient2Vec: A Personalized Interpretable Deep Representation of the Longitudinal Electronic Health Record
- Mixing Dirichlet Topic Models and Word Embeddings to Make lda2vec
- Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog Systems
- Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context
- Predicting the Popularity of Micro-videos with Multimodal Variational Encoder-Decoder Framework
- Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
- Generating Natural Language Explanations for Visual Question Answering using Scene Graphs and Visual Attention
- Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- Neural Net Models for Open-Domain Discourse Coherence
- Self-Gated Memory Recurrent Network for Efficient Scalable HDR Deghosting
- What comes next? Extractive summarization by next-sentence prediction
- Neural Contextual Conversation Learning with Labeled Question-Answering Pairs
- Review of state-of-the-arts in artificial intelligence with application to AI safety problem
- Instance-aware Image and Sentence Matching with Selective Multimodal LSTM
- Storytelling of Photo Stream with Bidirectional Multi-thread Recurrent Neural Network
- Deep Automated Multi-task Learning
- CAPS: Context Aware Personalized POI Sequence Recommender System
- soc2seq: Social Embedding meets Conversation Model
- Automatic coding of students' writing via Contrastive Representation Learning in the Wasserstein space
- Instance-based Transfer Learning for Multilingual Deep Retrieval
- Improving Context Aware Language Models
- Cross-modal Learning for Multi-modal Video Categorization
- Comparative Analysis of the Hidden Markov Model and LSTM: A Simulative Approach
- Low-Rank RNN Adaptation for Context-Aware Language Modeling
- EmotionX-DLC: Self-Attentive BiLSTM for Detecting Sequential Emotions in Dialogue
- Behavior Gated Language Models
- Estimating Fund-Raising Performance for Start-up Projects from a Market Graph Perspective
- Exploiting Temporal Coherence for Multi-modal Video Categorization
- Tweets Can Tell: Activity Recognition using Hybrid Long Short-Term Memory Model
- On-The-Fly Information Retrieval Augmentation for Language Models
- Knowledge Efficient Deep Learning for Natural Language Processing
- Estimating Early Fundraising Performance of Innovations via Graph-based Market Environment Model
- Modeling Dyadic Conversations for Personality Inference
- Image to Language Understanding: Captioning approach
- SAM: Semantic Attribute Modulation for Language Modeling and Style Variation
- LaNet: Real-time Lane Identification by Learning Road SurfaceCharacteristics from Accelerometer Data