Hierarchical Memory Networks
arXiv:1605.07427
Abstract
Memory networks are neural networks with an explicit memory component that can be both read and written to by the network. The memory is often addressed in a soft way using a softmax function, making end-to-end training with backpropagation possible. However, this is not computationally scalable for applications which require the network to read from extremely large memories. On the other hand, it is well known that hard attention mechanisms based on reinforcement learning are challenging to train successfully. In this paper, we explore a form of hierarchical memory network, which can be considered as a hybrid between hard and soft attention memory networks. The memory is organized in a hierarchical structure such that reading from it is done with less computation than soft attention over a flat memory, while also being easier to train than hard attention over a flat memory. Specifically, we propose to incorporate Maximum Inner Product Search (MIPS) in the training and inference procedures for our hierarchical memory network. We explore the use of various state-of-the art approximate MIPS techniques and report results on SimpleQuestions, a challenging large scale factoid question answering task.
10 pages
References in corpus (9)
- Adam: A Method for Stochastic Optimization
- Efficient Estimation of Word Representations in Vector Space
- Dynamic Memory Networks for Visual and Textual Question Answering
- Ask Me Anything: Dynamic Memory Networks for Natural Language Processing
- Large-scale Simple Question Answering with Memory Networks
- Asymmetric LSH (ALSH) for Sublinear Time Maximum Inner Product Search (MIPS)
- Evaluating Prerequisite Qualities for Learning End-to-End Dialog Systems
- Deep Networks With Large Output Spaces
- Clustering is Efficient for Approximate Maximum Inner Product Search
Cited by in corpus (26)
- Reformer: The Efficient Transformer
- Learning to Remember Rare Events
- Memory Matching Networks for One-Shot Image Recognition
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- MemexQA: Visual Memex Question Answering
- Augmenting Transformers with KNN-Based Composite Memory for Dialogue
- Question Answering from Unstructured Text by Retrieval and Comprehension
- Visual Question Answering with Memory-Augmented Networks
- A Read-Write Memory Network for Movie Story Understanding
- Multi-Hop Paragraph Retrieval for Open-Domain Question Answering
- Explore, Propose, and Assemble: An Interpretable Model for Multi-Hop Reading Comprehension
- Hierarchical Memory Networks for Answer Selection on Unknown Words
- Memorizing Comprehensively to Learn Adaptively: Unsupervised Cross-Domain Person Re-ID with Multi-level Memory
- Understanding and Improving Proximity Graph based Maximum Inner Product Search
- Adaptive Memory Networks
- Neural Machine Translation: A Review and Survey
- Single-View 3D Object Reconstruction from Shape Priors in Memory
- Knowledge Efficient Deep Learning for Natural Language Processing
- On-The-Fly Information Retrieval Augmentation for Language Models
- Contextual Memory Trees
- Energy-Efficient Inference Accelerator for Memory-Augmented Neural Networks on an FPGA
- -former: Infinite Memory Transformer
- Large Product Key Memory for Pretrained Language Models
- Memory networks for consumer protection:unfairness exposed
- Pyramid: A General Framework for Distributed Similarity Search