Explicit Knowledge-based Reasoning for Visual Question Answering
arXiv:1511.02570
Abstract
We describe a method for visual question answering which is capable of reasoning about contents of an image on the basis of information extracted from a large-scale knowledge base. The method not only answers natural language questions using concepts not contained in the image, but can provide an explanation of the reasoning by which it developed its answer. The method is capable of answering far more complex questions than the predominant long short-term memory-based approach, and outperforms it significantly in the testing. We also provide a dataset and a protocol by which to evaluate such methods, thus addressing one of the key issues in general visual ques- tion answering.
20 pages
References in corpus (4)
Cited by in corpus (17)
- Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- Zero-Shot Visual Question Answering
- Multi-modal Deep Analysis for Multimedia
- Visual Question Answering: A Survey of Methods and Datasets
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- FVQA: Fact-based Visual Question Answering
- Learning Visual Knowledge Memory Networks for Visual Question Answering
- OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
- An Empirical Evaluation of Visual Question Answering for Novel Objects
- BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge
- Understand, Compose and Respond - Answering Visual Questions by a Composition of Abstract Procedures
- Adversarial Multimodal Network for Movie Question Answering
- Adapting Visual Question Answering Models for Enhancing Multimodal Community Q&A Platforms