Yin and Yang: Balancing and Answering Binary Visual Questions
arXiv:1511.05099
Abstract
The complex compositional structure of language makes problems at the intersection of vision and language challenging. But language also provides a strong prior that can result in good superficial performance, without the underlying models truly understanding the visual content. This can hinder progress in pushing state of art in the computer vision aspects of multi-modal AI. In this paper, we address binary Visual Question Answering (VQA) on abstract scenes. We formulate this problem as visual verification of concepts inquired in the questions. Specifically, we convert the question to a tuple that concisely summarizes the visual concept to be detected in the image. If the concept can be found in the image, the answer to the question is "yes", and otherwise "no". Abstract scenes play two roles (1) They allow us to focus on the high-level semantics of the VQA task as opposed to the low-level recognition problems, and perhaps more importantly, (2) They provide us the modality to balance the dataset such that language priors are controlled, and the role of vision is essential. In particular, we collect fine-grained pairs of scenes for every question, such that the answer to the question is "yes" for one scene, and "no" for the other for the exact same question. Indeed, language priors alone do not perform better than chance on our balanced dataset. Moreover, our proposed approach matches the performance of a state-of-the-art VQA approach on the unbalanced dataset, and outperforms it on the balanced dataset.
References in corpus (10)
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Explain Images with Multimodal Recurrent Neural Networks
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Show and Tell: A Neural Image Caption Generator
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Exploring Nearest Neighbor Approaches for Image Captioning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Don't Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks
Cited by in corpus (13)
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- VQA: Visual Question Answering
- Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
- Zero-Shot Visual Question Answering
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Multi-modality Latent Interaction Network for Visual Question Answering
- An Analysis of Visual Question Answering Algorithms
- Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text
- VQA-LOL: Visual Question Answering under the Lens of Logic
- Learning to Disambiguate by Asking Discriminative Questions
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Transformation Driven Visual Reasoning