End-To-End Memory Networks
arXiv:1503.08895
Abstract
We introduce a neural network with a recurrent attention model over a possibly large external memory. The architecture is a form of Memory Network (Weston et al., 2015) but unlike the model in that work, it is trained end-to-end, and hence requires significantly less supervision during training, making it more generally applicable in realistic settings. It can also be seen as an extension of RNNsearch to the case where multiple computational steps (hops) are performed per output symbol. The flexibility of the model allows us to apply it to tasks as diverse as (synthetic) question answering and to language modeling. For the former our approach is competitive with Memory Networks, but with less supervision. For the latter, on the Penn TreeBank and Text8 datasets our approach demonstrates comparable performance to RNNs and LSTMs. In both cases we show that the key concept of multiple computational hops yields improved results.
Accepted to NIPS 2015
References in corpus (8)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Recurrent Neural Network Regularization
- DRAW: A Recurrent Neural Network For Image Generation
- Learning Longer Memory in Recurrent Neural Networks
- Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
- Neural Turing Machines
- Towards Neural Network-based Reasoning
Cited by in corpus (228)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Generalizing from a Few Examples: A Survey on Few-Shot Learning
- Attention in Natural Language Processing
- REALM: Retrieval-Augmented Language Model Pre-Training
- Generating Long Sequences with Sparse Transformers
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
- Multi-Task Learning with Deep Neural Networks: A Survey
- Go for a Walk and Arrive at the Answer: Reasoning Over Paths in Knowledge Bases using Reinforcement Learning
- Deep Learning Based Text Classification: A Comprehensive Review
- Universal Transformers
- Advances and Challenges in Conversational Recommender Systems: A Survey
- Improving Graph Neural Network Expressivity via Subgraph Isomorphism Counting
- Learning Multiagent Communication with Backpropagation
- Multimodal Residual Learning for Visual QA
- Neural Paraphrase Generation with Stacked Residual LSTM Networks
- Adversarial Text-to-Image Synthesis: A Review
- Natural Language Processing Advancements By Deep Learning: A Survey
- Fathom: Reference Workloads for Modern Deep Learning Methods
- Control of Memory, Active Perception, and Action in Minecraft
- Lifelong Sequential Modeling with Personalized Memorization for User Response Prediction
- Neural SLAM: Learning to Explore with External Memory
- A Neural Knowledge Language Model
- Associative Long Short-Term Memory
- Deep Learning for Technical Document Classification
- Hopfield Networks is All You Need
- Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
- Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose
- Graph Neural Networks for Natural Language Processing: A Survey
- Abstractive Summarization of Reddit Posts with Multi-level Memory Networks
- Recent Advances in Natural Language Inference: A Survey of Benchmarks, Resources, and Approaches
- Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes
- On the Binding Problem in Artificial Neural Networks
- Coarse-grain Fine-grain Coattention Network for Multi-evidence Question Answering
- Augmenting Self-attention with Persistent Memory
- DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis
- PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking
- PaperRobot: Incremental Draft Generation of Scientific Ideas
- MazeBase: A Sandbox for Learning from Games
- Compositional generalization through meta sequence-to-sequence learning
- Modifying Memories in Transformer Models
- BiERU: Bidirectional Emotional Recurrent Unit for Conversational Sentiment Analysis
- Learning from Few Samples: A Survey
- Temporal Self-Attention Network for Medical Concept Embedding
- PullNet: Open Domain Question Answering with Iterative Retrieval on Knowledge Bases and Text
- Global-to-local Memory Pointer Networks for Task-Oriented Dialogue
- A Taxonomy for Neural Memory Networks
- Deep Conversational Recommender in Travel
- Inductive Representation Learning on Temporal Graphs
- Uncertainty-Aware Attention for Reliable Interpretation and Prediction
- Addressing Some Limitations of Transformers with Feedback Memory
- Compositional Generalization by Learning Analytical Expressions
- Neural Machine Reading Comprehension: Methods and Trends
- Signed Distance-based Deep Memory Recommender
- Simulating Action Dynamics with Neural Process Networks
- Metalearned Neural Memory
- Towards Interpretable Reinforcement Learning Using Attention Augmented Agents
- Generating Radiology Reports via Memory-driven Transformer
- PonderNet: Learning to Ponder
- Learning Video Object Segmentation from Unlabeled Videos
- Progressive Memory Banks for Incremental Domain Adaptation
- Augmenting Transformers with KNN-Based Composite Memory for Dialogue
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- Learning to Remember Patterns: Pattern Matching Memory Networks for Traffic Forecasting
- Memory-augmented Attention Modelling for Videos
- Self-Attentive Residual Decoder for Neural Machine Translation
- Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading
- Sparse and Continuous Attention Mechanisms
- CmnRec: Sequential Recommendations with Chunk-accelerated Memory Network
- Object Files and Schemata: Factorizing Declarative and Procedural Knowledge in Dynamical Systems
- Learning in Text Streams: Discovery and Disambiguation of Entity and Relation Instances
- Conditional Self-Attention for Query-based Summarization
- Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation
- Learning End-to-End Goal-Oriented Dialog with Maximal User Task Success and Minimal Human Agent Use
- Graph Message Passing with Cross-location Attentions for Long-term ILI Prediction
- Commonsense for Generative Multi-Hop Question Answering Tasks
- A Short Survey On Memory Based Reinforcement Learning
- Real-Time Emotion Recognition via Attention Gated Hierarchical Memory Network
- Joint Modeling of Local and Global Temporal Dynamics for Multivariate Time Series Forecasting with Missing Values
- Progressive Attention Memory Network for Movie Story Question Answering
- Neural Assistant: Joint Action Prediction, Response Generation, and Latent Knowledge Reasoning
- VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop
- DialogueCRN: Contextual Reasoning Networks for Emotion Recognition in Conversations
- Knowledge-aware Attention Network for Protein-Protein Interaction Extraction
- Leveraging Prior Knowledge for Protein-Protein Interaction Extraction with Memory Network
- GNN is a Counter? Revisiting GNN for Question Answering
- Kernelized Memory Network for Video Object Segmentation
- Improving Person Re-identification with Iterative Impression Aggregation
- Exploiting Human Social Cognition for the Detection of Fake and Fraudulent Faces via Memory Networks
- Memory-Based Graph Networks
- OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
- Graph-based Multi-hop Reasoning for Long Text Generation
- PREMIER: Personalized REcommendation for Medical prescrIptions from Electronic Records
- Towards Persona-Based Empathetic Conversational Models
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QA
- Multi-Hop Paragraph Retrieval for Open-Domain Question Answering
- Improving Conditioning in Context-Aware Sequence to Sequence Models
- Machine Learning with World Knowledge: The Position and Survey
- Deep Learning for Multi-Scale Changepoint Detection in Multivariate Time Series
- Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented Dialog
- Mutual Information Scaling and Expressive Power of Sequence Models
- Learning to Respond with Your Favorite Stickers: A Framework of Unifying Multi-Modality and User Preference in Multi-Turn Dialog
- Differentiable Logic Machines
- Transfer Meets Hybrid: A Synthetic Approach for Cross-Domain Collaborative Filtering with Text
- Learning Deep Matrix Representations
- A Mathematical Theory of Attention
- Advances in Natural Language Question Answering: A Review
- Domain Adaptation with Auxiliary Target Domain-Oriented Classifier
- Modeling relation paths for knowledge base completion via joint adversarial training
- The Variational Bandwidth Bottleneck: Stochastic Evaluation on an Information Budget
- A Multi-Type Multi-Span Network for Reading Comprehension that Requires Discrete Reasoning
- MEMO: A Deep Network for Flexible Combination of Episodic Memories
- Entity-Consistent End-to-end Task-Oriented Dialogue System with KB Retriever
- Meaningful Answer Generation of E-Commerce Question-Answering
- CPGAN: Full-Spectrum Content-Parsing Generative Adversarial Networks for Text-to-Image Synthesis
- Differentiable programming and its applications to dynamical systems
- Recovering Dropped Pronouns in Chinese Conversations via Modeling Their Referents
- Long Short-Term Attention
- Integrative Analysis of Patient Health Records and Neuroimages via Memory-based Graph Convolutional Network
- Learning to learn with backpropagation of Hebbian plasticity
- Enriching a Model's Notion of Belief using a Persistent Memory
- Untangling tradeoffs between recurrence and self-attention in neural networks
- Multi-Task Learning for Conversational Question Answering over a Large-Scale Knowledge Base
- Compositional Generalization with Tree Stack Memory Units
- Identifying Sub-Phenotypes of Acute Kidney Injury using Structured and Unstructured Electronic Health Record Data with Memory Networks
- Multi-source Attention for Unsupervised Domain Adaptation
- A Framework for Evaluation of Machine Reading Comprehension Gold Standards
- SGoLAM: Simultaneous Goal Localization and Mapping for Multi-Object Goal Navigation
- Product Kanerva Machines: Factorized Bayesian Memory
- Be Concise and Precise: Synthesizing Open-Domain Entity Descriptions from Facts
- Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
- Episodic Memory Reader: Learning What to Remember for Question Answering from Streaming Data
- Video Imprint
- Unsupervised Learning of KB Queries in Task-Oriented Dialogs
- Neural Status Registers
- Explicit-Blurred Memory Network for Analyzing Patient Electronic Health Records
- Integrating Dictionary Feature into A Deep Learning Model for Disease Named Entity Recognition
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- Scalable and Accurate Dialogue State Tracking via Hierarchical Sequence Generation
- DSReg: Using Distant Supervision as a Regularizer
- Staircase Attention for Recurrent Processing of Sequences
- Task-Oriented Conversation Generation Using Heterogeneous Memory Networks
- GraphDialog: Integrating Graph Knowledge into End-to-End Task-Oriented Dialogue Systems
- DAWN: Dual Augmented Memory Network for Unsupervised Video Object Tracking
- Ain't Nobody Got Time For Coding: Structure-Aware Program Synthesis From Natural Language
- Memory-Based Neighbourhood Embedding for Visual Recognition
- Visual Tracking via Dynamic Memory Networks
- Multilingual Dialogue Generation with Shared-Private Memory
- Neural Stored-program Memory
- On the Regularity of Attention
- The Tensor Brain: A Unified Theory of Perception, Memory and Semantic Decoding
- Towards a Universal Continuous Knowledge Base
- Learning Feature Aggregation for Deep 3D Morphable Models
- Why Do We Click: Visual Impression-aware News Recommendation
- Understanding and Controlling Memory in Recurrent Neural Networks
- Single-View 3D Object Reconstruction from Shape Priors in Memory
- Trying Bilinear Pooling in Video-QA
- Verifying Security Protocols using Dynamic Strategies
- Neural Machine Translation: A Review and Survey
- Microblog Hashtag Generation via Encoding Conversation Contexts
- Incremental Concept Learning via Online Generative Memory Recall
- Learning Associative Inference Using Fast Weight Memory
- Language Model Evaluation in Open-ended Text Generation
- Question-Aware Memory Network for Multi-hop Question Answering in Human-Robot Interaction
- Sparse Continuous Distributions and Fenchel-Young Losses
- Learning Algorithmic Solutions to Symbolic Planning Tasks with a Neural Computer Architecture
- Memory and attention in deep learning
- Questions to Guide the Future of Artificial Intelligence Research
- Recursive Sketches for Modular Deep Learning
- Better Long-Range Dependency By Bootstrapping A Mutual Information Regularizer
- Reconciling the Discrete-Continuous Divide: Towards a Mathematical Theory of Sparse Communication
- End-to-End Egospheric Spatial Memory
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- Explainable Deep RDFS Reasoner
- Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
- pix2rule: End-to-end Neuro-symbolic Rule Learning
- Deep Visual Odometry with Adaptive Memory
- Adaptive Knowledge-Enhanced Bayesian Meta-Learning for Few-shot Event Detection
- Relational dynamic memory networks
- Memory-Augmented Temporal Dynamic Learning for Action Recognition
- Contrastive Language Adaptation for Cross-Lingual Stance Detection
- Explainable Neural Computation via Stack Neural Module Networks
- 3D Meta Point Signature: Learning to Learn 3D Point Signature for 3D Dense Shape Correspondence
- DNC-Aided SCL-Flip Decoding of Polar Codes
- RotLSTM: Rotating Memories in Recurrent Neural Networks
- A Proposal for Intelligent Agents with Episodic Memory
- Neural Program Meta-Induction
- Complex Knowledge Base Question Answering: A Survey
- Sequential Recommender via Time-aware Attentive Memory Network
- Prototype Matching Networks for Large-Scale Multi-label Genomic Sequence Classification
- Automatically Exposing Problems with Neural Dialog Models
- Towards Continual Entity Learning in Language Models for Conversational Agents
- DMV: Visual Object Tracking via Part-level Dense Memory and Voting-based Retrieval
- TME-BNA: Temporal Motif-Preserving Network Embedding with Bicomponent Neighbor Aggregation
- Enhancing Reinforcement Learning with discrete interfaces to learn the Dyck Language
- Learning Representations for Zero-Shot Retrieval over Structured Data
- Training with Streaming Annotation
- AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue
- On Estimating the Training Cost of Conversational Recommendation Systems
- Video SemNet: Memory-Augmented Video Semantic Network
- Unsupervised Keyword Extraction for Full-sentence VQA
- Speaker-Sensitive Dual Memory Networks for Multi-Turn Slot Tagging
- Document-level Neural Machine Translation with Associated Memory Network
- Data-Efficient Methods for Dialogue Systems
- Learning Invariants through Soft Unification
- CNN with large memory layers
- Recoding latent sentence representations -- Dynamic gradient-based activation modification in RNNs
- End-to-End Video Question-Answer Generation with Generator-Pretester Network
- Probabilistic Metric Learning with Adaptive Margin for Top-K Recommendation
- Learning from Web Data with Self-Organizing Memory Module
- Towards conceptual generalization in the embedding space
- Adaptive Attention Span in Transformers
- BERT Embeddings Can Track Context in Conversational Search
- Aerial Scene Understanding in The Wild: Multi-Scene Recognition via Prototype-based Memory Networks
- Building A User-Centric and Content-Driven Socialbot
- Response-Anticipated Memory for On-Demand Knowledge Integration in Response Generation
- A Real-time Action Representation with Temporal Encoding and Deep Compression
- Memory networks for consumer protection:unfairness exposed
- Online Multi-modal Person Search in Videos
- Relation/Entity-Centric Reading Comprehension
- BiteNet: Bidirectional Temporal Encoder Network to Predict Medical Outcomes
- Why can't memory networks read effectively?
- Memory-Augmented Recurrent Networks for Dialogue Coherence
- NE-Table: A Neural key-value table for Named Entities
- Learning to Organize Knowledge and Answer Questions with N-Gram Machines
- Transferable End-to-End Aspect-based Sentiment Analysis with Selective Adversarial Learning
- Survey of reasoning using Neural networks
- What do Entity-Centric Models Learn? Insights from Entity Linking in Multi-Party Dialogue
- Generalizable and Explainable Dialogue Generation via Explicit Action Learning