Neural Self Talk: Image Understanding via Continuous Questioning and Answering
arXiv:1512.03460
Abstract
In this paper we consider the problem of continuously discovering image contents by actively asking image based questions and subsequently answering the questions being asked. The key components include a Visual Question Generation (VQG) module and a Visual Question Answering module, in which Recurrent Neural Networks (RNN) and Convolutional Neural Network (CNN) are used. Given a dataset that contains images, questions and their answers, both modules are trained at the same time, with the difference being VQG uses the images as input and the corresponding questions as output, while VQA uses images and questions as input and the corresponding answers as output. We evaluate the self talk process subjectively using Amazon Mechanical Turk, which show effectiveness of the proposed method.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Explain Images with Multimodal Recurrent Neural Networks
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Learning a Recurrent Visual Representation for Image Caption Generation
- CIDEr: Consensus-based Image Description Evaluation
- From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge