Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
arXiv:1511.06973
Abstract
We propose a method for visual question answering which combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions. This allows more complex questions to be answered using the predominant neural network-based approach than has previously been possible. It particularly allows questions to be asked about the contents of an image, even when the image itself does not contain the whole answer. The method constructs a textual representation of the semantic content of an image, and merges it with textual information sourced from a knowledge base, to develop a deeper understanding of the scene viewed. Priming a recurrent neural network with this combined information, and the submitted question, leads to a very flexible visual question answering approach. We are specifically able to answer questions posed in natural language, that refer to information not contained in the image. We demonstrate the effectiveness of our model on two publicly available datasets, Toronto COCO-QA and MS COCO-VQA and show that it produces the best reported results in both cases.
Accepted to IEEE Conf. Computer Vision and Pattern Recognition
References in corpus (16)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Microsoft COCO Captions: Data Collection and Evaluation Server
- VQA: Visual Question Answering
- CNN: Single-label to Multi-label
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Describing Videos by Exploiting Temporal Structure
- Show and Tell: A Neural Image Caption Generator
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Learning a Recurrent Visual Representation for Image Caption Generation
- Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
- What value do explicit high level concepts have in vision to language problems?
- Don't Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks
Cited by in corpus (19)
- Simple Baseline for Visual Question Answering
- A Focused Dynamic Attention Model for Visual Question Answering
- Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- Zero-shot Recognition via Semantic Embeddings and Knowledge Graphs
- Task-driven Visual Saliency and Attention-based Visual Question Answering
- An Analysis of Visual Question Answering Algorithms
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Asking the Difficult Questions: Goal-Oriented Visual Question Generation via Intermediate Rewards
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Video Fill in the Blank with Merging LSTMs
- Learning Visual Knowledge Memory Networks for Visual Question Answering
- Image-Question-Answer Synergistic Network for Visual Dialog
- An Empirical Evaluation of Visual Question Answering for Novel Objects
- MarioQA: Answering Questions by Watching Gameplay Videos
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Long Activity Video Understanding using Functional Object-Oriented Network