Visual Question Answering: Datasets, Algorithms, and Future Challenges
arXiv:1610.01465 · doi:10.1016/j.cviu.2017.06.005
Abstract
Visual Question Answering (VQA) is a recent problem in computer vision and natural language processing that has garnered a large amount of interest from the deep learning, computer vision, and natural language processing communities. In VQA, an algorithm needs to answer text-based questions about images. Since the release of the first VQA dataset in 2014, additional datasets have been released and many algorithms have been proposed. In this review, we critically examine the current state of VQA in terms of problem formulation, existing datasets, evaluation metrics, and algorithms. In particular, we discuss the limitations of current datasets with regard to their ability to properly train and assess VQA algorithms. We then exhaustively review existing algorithms for VQA. Finally, we discuss possible future directions for VQA and image understanding research.
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Fully Convolutional Networks for Semantic Segmentation
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- Multiple Object Recognition with Visual Attention
- Dynamic Memory Networks for Visual and Textual Question Answering
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Multimodal Residual Learning for Visual QA
- Hadamard Product for Low-rank Bilinear Pooling
- Exploring Nearest Neighbor Approaches for Image Captioning
- A Focused Dynamic Attention Model for Visual Question Answering
- Training Recurrent Answering Units with Joint Loss Minimization for VQA
- DualNet: Domain-Invariant Network for Visual Question Answering
Cited by in corpus (39)
- Unleashing the potential of prompt engineering for large language models
- VLP: A Survey on Vision-Language Pre-training
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- Bangla Natural Language Processing: A Comprehensive Analysis of Classical, Machine Learning, and Deep Learning Based Methods
- A Question-Centric Model for Visual Question Answering in Medical Imaging
- Recent Advances in Natural Language Inference: A Survey of Benchmarks, Resources, and Approaches
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- DVQA: Understanding Data Visualizations via Question Answering
- Structured Attentions for Visual Question Answering
- An Analysis of Visual Question Answering Algorithms
- A Review of Emerging Research Directions in Abstract Visual Reasoning
- REMIND Your Neural Network to Prevent Catastrophic Forgetting
- Vision Skills Needed to Answer Visual Questions
- Learning Convolutional Text Representations for Visual Question Answering
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering
- Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering
- CoLLIE: Continual Learning of Language Grounding from Language-Image Embeddings
- Exposing and Correcting the Gender Bias in Image Captioning Datasets and Models
- Figure Captioning with Reasoning and Sequence-Level Training
- Semi-Supervised Panoptic Narrative Grounding
- Answer Them All! Toward Universal Visual Question Answering Models
- Zero-Shot Transfer VQA Dataset
- EKTVQA: Generalized use of External Knowledge to empower Scene Text in Text-VQA
- TallyQA: Answering Complex Counting Questions
- Answering Questions about Data Visualizations using Efficient Bimodal Fusion
- Visual Question Answering with Prior Class Semantics
- Quantifying and Alleviating the Language Prior Problem in Visual Question Answering
- New Ideas and Trends in Deep Multimodal Content Understanding: A Review
- Challenges and Prospects in Vision and Language Research
- Learning to Localize Sound Sources in Visual Scenes: Analysis and Applications
- Improved RAMEN: Towards Domain Generalization for Visual Question Answering
- Understand, Compose and Respond - Answering Visual Questions by a Composition of Abstract Procedures
- The Wisdom of MaSSeS: Majority, Subjectivity, and Semantic Similarity in the Evaluation of VQA
- Selective Replay Enhances Learning in Online Continual Analogical Reasoning
- On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering
- Customized Image Narrative Generation via Interactive Visual Question Generation and Answering
- Understanding in Artificial Intelligence
- Empirically Verifying Hypotheses Using Reinforcement Learning