Long Short-Term Memory-Networks for Machine Reading
arXiv:1601.06733
Abstract
In this paper we address the question of how to render sequence-level networks better at handling structured input. We propose a machine reading simulator which processes text incrementally from left to right and performs shallow reasoning with memory and attention. The reader extends the Long Short-Term Memory architecture with a memory network in place of a single memory cell. This enables adaptive memory usage during recurrence with neural attention, offering a way to weakly induce relations among tokens. The system is initially designed to process a single sequence but we also demonstrate how to integrate it with an encoder-decoder architecture. Experiments on language modeling, sentiment analysis, and natural language inference show that our model matches or outperforms the state of the art.
Published as a conference paper at EMNLP 2016
References in corpus (9)
- On the difficulty of training Recurrent Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Convolutional Neural Networks for Sentence Classification
- Dynamic Memory Networks for Visual and Textual Question Answering
- Transition-Based Dependency Parsing with Stack Long Short-Term Memory
- A Convolutional Neural Network for Modelling Sentences
- Learning Longer Memory in Recurrent Neural Networks
- Learning to Compose Neural Networks for Question Answering
- Recurrent Memory Networks for Language Modeling
Cited by in corpus (93)
- A Deep Reinforced Model for Abstractive Summarization
- Enhanced LSTM for Natural Language Inference
- Deep Learning Based Text Classification: A Comprehensive Review
- Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
- CCNet: Criss-Cross Attention for Semantic Segmentation
- Neural Natural Language Inference Models Enhanced with External Knowledge
- OpenTag: Open Attribute Value Extraction from Product Profiles [Deep Learning, Active Learning, Named Entity Recognition]
- Bilateral Multi-Perspective Matching for Natural Language Sentences
- ATRank: An Attention-Based User Behavior Modeling Framework for Recommendation
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- A Survey on Text Classification: From Shallow to Deep Learning
- Global-Locally Self-Attentive Dialogue State Tracker
- CLVSA: A Convolutional LSTM Based Variational Sequence-to-Sequence Model with Attention for Predicting Trends of Financial Markets
- Dynamic Integration of Background Knowledge in Neural NLU Systems
- A Decomposable Attention Model for Natural Language Inference
- Reinforced Mnemonic Reader for Machine Reading Comprehension
- Deep Learning Based Chatbot Models
- Graph-based Knowledge Distillation by Multi-head Attention Network
- Deep Implicit Coordination Graphs for Multi-agent Reinforcement Learning
- A Fast Unified Model for Parsing and Sentence Understanding
- Attentive Group Equivariant Convolutional Networks
- Memory-enhanced Decoder for Neural Machine Translation
- Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor Networks
- Neural Language Modeling by Jointly Learning Syntax and Lexicon
- Deriving Neural Architectures from Sequence and Graph Kernels
- Textual Entailment with Structured Attentions and Composition
- Attention-Based Neural Networks for Chroma Intra Prediction in Video Coding
- PointGrow: Autoregressively Learned Point Cloud Generation with Self-Attention
- Learning to Focus: Cascaded Feature Matching Network for Few-shot Image Recognition
- Co-Attentive Equivariant Neural Networks: Focusing Equivariance On Transformations Co-Occurring In Data
- Separating Answers from Queries for Neural Reading Comprehension
- Generating Natural Language Inference Chains
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- Aspect-augmented Adversarial Networks for Domain Adaptation
- Building a Neural Semantic Parser from a Domain Ontology
- Transformer Based Reinforcement Learning For Games
- Enhancing Machine Translation with Dependency-Aware Self-Attention
- Natural Language Generation with Neural Variational Models
- Learning Structured Text Representations
- A Neural Architecture Mimicking Humans End-to-End for Natural Language Inference
- Why an Android App is Classified as Malware? Towards Malware Classification Interpretation
- On the Effective Use of Pretraining for Natural Language Inference
- Transformers for One-Shot Visual Imitation
- When FastText Pays Attention: Efficient Estimation of Word Representations using Constrained Positional Weighting
- Multi-task Learning over Graph Structures
- Long Short-Term Attention
- Point Clouds Learning with Attention-based Graph Convolution Networks
- MQTransformer: Multi-Horizon Forecasts with Context Dependent and Feedback-Aware Attention
- Jointly Trained Sequential Labeling and Classification by Sparse Attention Neural Networks
- Modeling Document Interactions for Learning to Rank with Regularized Self-Attention
- An Empirical Evaluation of various Deep Learning Architectures for Bi-Sequence Classification Tasks
- Second-Order Word Embeddings from Nearest Neighbor Topological Features
- Multitask Learning for Class-Imbalanced Discourse Classification
- Attention-based Memory Selection Recurrent Network for Language Modeling
- FineText: Text Classification via Attention-based Language Model Fine-tuning
- AttaNet: Attention-Augmented Network for Fast and Accurate Scene Parsing
- Group Equivariant Stand-Alone Self-Attention For Vision
- Attention Boosted Sequential Inference Model
- Chroma Intra Prediction with attention-based CNN architectures
- Attentive Convolution: Equipping CNNs with RNN-style Attention Mechanisms
- Endowing Deep 3D Models with Rotation Invariance Based on Principal Component Analysis
- Why and How to Pay Different Attention to Phrase Alignments of Different Intensities
- Deep Fourier Kernel for Self-Attentive Point Processes
- What If We Simply Swap the Two Text Fragments? A Straightforward yet Effective Way to Test the Robustness of Methods to Confounding Signals in Nature Language Inference Tasks
- Sentence Encoding with Tree-constrained Relation Networks
- Syntax-based Attention Model for Natural Language Inference
- Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks
- A Feasible Framework for Arbitrary-Shaped Scene Text Recognition
- Malware Classification Using Long Short-Term Memory Models
- Look-ahead Attention for Generation in Neural Machine Translation
- Learning an Executable Neural Semantic Parser
- Relevance-Promoting Language Model for Short-Text Conversation
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- VTAMIQ: Transformers for Attention Modulated Image Quality Assessment
- Exploiting Inter-pixel Correlations in Unsupervised Domain Adaptation for Semantic Segmentation
- End-Task Oriented Textual Entailment via Deep Explorations of Inter-Sentence Interactions
- Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance Approach
- Prototypical Recurrent Unit
- Less Memory, Faster Speed: Refining Self-Attention Module for Image Reconstruction
- A memory enhanced LSTM for modeling complex temporal dependencies
- An Iterative Contextualization Algorithm with Second-Order Attention
- Middle-Out Decoding
- Incremental Scene Synthesis
- Exploiting Sentence Embedding for Medical Question Answering
- Using Context Information to Enhance Simple Question Answering
- Ensemble ALBERT on SQuAD 2.0
- Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences
- Improving Neural Language Models by Segmenting, Attending, and Predicting the Future
- Two-Stream Appearance Transfer Network for Person Image Generation
- Interpretable Structure-aware Document Encoders with Hierarchical Attention
- Emotion Detection with Neural Personal Discrimination
- Memory networks for consumer protection:unfairness exposed
- MuSLCAT: Multi-Scale Multi-Level Convolutional Attention Transformer for Discriminative Music Modeling on Raw Waveforms