Stacked Attention Networks for Image Question Answering
arXiv:1511.02274
Abstract
This paper presents stacked attention networks (SANs) that learn to answer natural language questions from images. SANs use semantic representation of a question as query to search for the regions in an image that are related to the answer. We argue that image question answering (QA) often requires multiple steps of reasoning. Thus, we develop a multiple-layer SAN in which we query an image multiple times to infer the answer progressively. Experiments conducted on four image QA data sets demonstrate that the proposed SANs significantly outperform previous state-of-the-art approaches. The visualization of the attention layers illustrates the progress that the SAN locates the relevant visual clues that lead to the answer of the question layer-by-layer.
test-dev/standard results added
Cited by in corpus (53)
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Recent Advances in Deep Learning: An Overview
- Learning to Reason: End-to-End Module Networks for Visual Question Answering
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Large-Scale Image Retrieval with Attentive Deep Local Features
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Dual Attention Networks for Multimodal Reasoning and Matching
- Revisiting Video Saliency: A Large-scale Benchmark and a New Model
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Multi-scale Deep Learning Architectures for Person Re-identification
- Creativity: Generating Diverse Questions using Variational Autoencoders
- It Takes Two to Tango: Towards Theory of AI's Mind
- Iterative Visual Reasoning Beyond Convolutions
- Less Is More: Picking Informative Frames for Video Captioning
- DDRprog: A CLEVR Differentiable Dynamic Reasoning Programmer
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Task-driven Visual Saliency and Attention-based Visual Question Answering
- An Analysis of Visual Question Answering Algorithms
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- GuessWhat?! Visual object discovery through multi-modal dialogue
- Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System
- Overcoming Language Priors in Visual Question Answering with Adversarial Regularization
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Learning Visual Knowledge Memory Networks for Visual Question Answering
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- Compact Global Descriptor for Neural Networks
- Explainable High-order Visual Question Reasoning: A New Benchmark and Knowledge-routed Network
- Question Type Guided Attention in Visual Question Answering
- Dual-Glance Model for Deciphering Social Relationships
- MarioQA: Answering Questions by Watching Gameplay Videos
- Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation
- On Attention Modules for Audio-Visual Synchronization
- Mean Box Pooling: A Rich Image Representation and Output Embedding for the Visual Madlibs Task
- Learning to Disambiguate by Asking Discriminative Questions
- Embodied Language Grounding with 3D Visual Feature Representations
- Spatial Memory for Context Reasoning in Object Detection
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- Imitation Learning of Robot Policies by Combining Language, Vision and Demonstration
- When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in Vlogs
- Transfer Learning in Visual and Relational Reasoning
- Pay Attention to Those Sets! Learning Quantification from Images
- Applying recent advances in Visual Question Answering to Record Linkage
- Timestamping Documents and Beliefs
- MOC-GAN: Mixing Objects and Captions to Generate Realistic Images
- Visual Question Answering Using Semantic Information from Image Descriptions
- Proposing Plausible Answers for Open-ended Visual Question Answering
- Learning Robust Video Synchronization without Annotations
- Customized Image Narrative Generation via Interactive Visual Question Generation and Answering
- Human-Aware Motion Deblurring