Visual7W: Grounded Question Answering in Images
arXiv:1511.03416
Abstract
We have seen great progress in basic perceptual tasks such as object recognition and detection. However, AI models still fail to match humans in high-level vision tasks due to the lack of capacities for deeper reasoning. Recently the new task of visual question answering (QA) has been proposed to evaluate a model's capacity for deep image understanding. Previous works have established a loose, global association between QA sentences and images. However, many questions and answers, in practice, relate to local regions in the images. We establish a semantic link between textual descriptions and image regions by object-level grounding. It enables a new type of QA with visual answers, in addition to textual answers used in previous work. We study the visual QA tasks in a grounded setting with a large collection of 7W multiple-choice QA pairs. Furthermore, we evaluate human performance and several baseline models on the QA tasks. Finally, we propose a novel LSTM model with spatial attention to tackle the 7W QA tasks.
CVPR 2016
References in corpus (11)
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- DeepPose: Human Pose Estimation via Deep Neural Networks
- VQA: Visual Question Answering
- DRAW: A Recurrent Neural Network For Image Generation
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
Cited by in corpus (22)
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Identity-Aware Textual-Visual Matching with Latent Co-attention
- Revisiting Visual Question Answering Baselines
- Context-aware Captions from Context-agnostic Supervision
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Task-driven Visual Saliency and Attention-based Visual Question Answering
- A Fast and Accurate One-Stage Approach to Visual Grounding
- An Analysis of Visual Question Answering Algorithms
- A Diagram Is Worth A Dozen Images
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- An Empirical Evaluation of Visual Question Answering for Novel Objects
- VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
- C3VQG: Category Consistent Cyclic Visual Question Generation
- Learning to Disambiguate by Asking Discriminative Questions
- Spatial Memory for Context Reasoning in Object Detection
- Point and Ask: Incorporating Pointing into Visual Question Answering
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- Recurrent Multimodal Interaction for Referring Image Segmentation
- Spatial Attention as an Interface for Image Captioning Models