Robustness Analysis of Visual QA Models by Basic Questions
arXiv:1709.04625
Abstract
Visual Question Answering (VQA) models should have both high robustness and accuracy. Unfortunately, most of the current VQA research only focuses on accuracy because there is a lack of proper methods to measure the robustness of VQA models. There are two main modules in our algorithm. Given a natural language question about an image, the first module takes the question as input and then outputs the ranked basic questions, with similarity scores, of the main given question. The second module takes the main question, image and these basic questions as input and then outputs the text-based answer of the main question about the given image. We claim that a robust VQA model is one, whose performance is not changed much when related basic questions as also made available to it as input. We formulate the basic questions generation problem as a LASSO optimization, and also propose a large scale Basic Question Dataset (BQD) and Rscore (novel robustness measure), for analyzing the robustness of VQA models. We hope our BQD will be used as a benchmark for to evaluate the robustness of VQA models, so as to help the community build more robust and accurate VQA models.
Accepted by CVPR 2018 VQA Challenge and Visual Dialog Workshop. (Acknowledgement updating)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- CIDEr: Consensus-based Image Description Evaluation
- VQABQ: Visual Question Answering by Basic Questions
Cited by in corpus (7)
- A Novel Hybrid Machine Learning Model for Auto-Classification of Retinal Diseases
- Contextualized Keyword Representations for Multi-modal Retinal Image Captioning
- Query-controllable Video Summarization
- DeepOpht: Medical Report Generation for Retinal Images via Deep Models and Visual Explanation
- Auto-Classification of Retinal Diseases in the Limit of Sparse Data Using a Two-Streams Machine Learning Model
- GPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video Summarization
- Robust Unsupervised Multi-Object Tracking in Noisy Environments