Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
arXiv:1505.05612
Abstract
In this paper, we present the mQA model, which is able to answer questions about the content of an image. The answer can be a sentence, a phrase or a single word. Our model contains four components: a Long Short-Term Memory (LSTM) to extract the question representation, a Convolutional Neural Network (CNN) to extract the visual representation, an LSTM for storing the linguistic context in an answer, and a fusing component to combine the information from the first three components and generate the answer. We construct a Freestyle Multilingual Image Question Answering (FM-IQA) dataset to train and evaluate our mQA model. It contains over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations. The quality of the generated answers of our mQA model on this dataset is evaluated by human judges through a Turing Test. Specifically, we mix the answers provided by humans and our model. The human judges need to distinguish our model from the human. They will also provide a score (i.e. 0, 1, 2, the larger the better) indicating the quality of the answer. We propose strategies to monitor the quality of this evaluation process. The experiments show that in 64.7% of cases, the human judges cannot distinguish our model from humans. The average score is 1.454 (1.918 for human). The details of this work, including the FM-IQA dataset, can be found on the project page: http://idl.baidu.com/FM-IQA.html
Dataset released on the project page, see http://idl.baidu.com/FM-IQA.html ; NIPS 2015 camera ready version
References in corpus (12)
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- VQA: Visual Question Answering
- Explain Images with Multimodal Recurrent Neural Networks
- Learning Longer Memory in Recurrent Neural Networks
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Learning a Recurrent Visual Representation for Image Caption Generation
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Simple Image Description Generator via a Linear Phrase-Based Approach
Cited by in corpus (62)
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- VQA: Visual Question Answering
- A simple neural network module for relational reasoning
- Exploring Models and Data for Image Question Answering
- Simple Baseline for Visual Question Answering
- ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering
- Visual Question Answering: Datasets, Algorithms, and Future Challenges
- Generation and Comprehension of Unambiguous Object Descriptions
- Fluency-Guided Cross-Lingual Image Captioning
- Explicit Knowledge-based Reasoning for Visual Question Answering
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
- Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Learning Deep Structure-Preserving Image-Text Embeddings
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- Vision-to-Language Tasks Based on Attributes and Attention Mechanism
- M6: A Chinese Multimodal Pretrainer
- Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
- Visual7W: Grounded Question Answering in Images
- Visual Question Answering: A Survey of Methods and Datasets
- Uncovering Temporal Context for Video Question and Answering
- Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Yin and Yang: Balancing and Answering Binary Visual Questions
- FVQA: Fact-based Visual Question Answering
- Compositional Memory for Visual Question Answering
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- What value do explicit high level concepts have in vision to language problems?
- Creativity: Generating Diverse Questions using Variational Autoencoders
- TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
- Revisiting Visual Question Answering Baselines
- An Analysis of Visual Question Answering Algorithms
- Multi-Cue Zero-Shot Learning with Strong Supervision
- Stories for Images-in-Sequence by using Visual and Narrative Components
- Neural Self Talk: Image Understanding via Continuous Questioning and Answering
- Visual Question Answering with Memory-Augmented Networks
- Pre-Trained Models: Past, Present and Future
- Factor Graph Attention
- CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication
- Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
- Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven's Progressive Matrices
- Adversarial Attacks Beyond the Image Space
- Learning Visual Question Answering by Bootstrapping Hard Attention
- Fooling Vision and Language Models Despite Localization and Attention Mechanism
- Understanding Image and Text Simultaneously: a Dual Vision-Language Machine Comprehension Task
- End-to-end Image Captioning Exploits Multimodal Distributional Similarity
- Mean Box Pooling: A Rich Image Representation and Output Embedding for the Visual Madlibs Task
- Spatial Memory for Context Reasoning in Object Detection
- Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
- Answering Image Riddles using Vision and Reasoning through Probabilistic Soft Logic
- Relationship-Embedded Representation Learning for Grounding Referring Expressions
- Visual Question Answering as Reading Comprehension
- Learning Models for Actions and Person-Object Interactions with Transfer to Question Answering
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- Designing Multimodal Datasets for NLP Challenges
- Leveraging Visual Question Answering for Image-Caption Ranking
- Towards Solving Multimodal Comprehension
- Pay Attention to Those Sets! Learning Quantification from Images