Interpretable Counting for Visual Question Answering
arXiv:1712.08697
Abstract
Questions that require counting a variety of objects in images remain a major challenge in visual question answering (VQA). The most common approaches to VQA involve either classifying answers based on fixed length representations of both the image and question or summing fractional counts estimated from each section of the image. In contrast, we treat counting as a sequential decision process and force our model to make discrete choices of what to count. Specifically, the model sequentially selects from detected objects and learns interactions between objects that influence subsequent selections. A distinction of our approach is its intuitive and interpretable output, as discrete counts are automatically grounded in the image. Furthermore, our method outperforms the state of the art architecture for VQA on multiple metrics that evaluate counting.
ICLR 2018
Cited by in corpus (13)
- Learning to Count Objects in Natural Images for Visual Question Answering
- Bilinear Attention Networks
- TVQA+: Spatio-Temporal Grounding for Video Question Answering
- Human-Adversarial Visual Question Answering
- In Defense of Grid Features for Visual Question Answering
- MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond
- Linguistically-aware Attention for Reducing the Semantic-Gap in Vision-Language Tasks
- Object-Centric Diagnosis of Visual Reasoning
- Interpretable Visual Question Answering by Reasoning on Dependency Trees
- TallyQA: Answering Complex Counting Questions
- Overcoming Statistical Shortcuts for Open-ended Visual Counting
- Toward Interpretability of Dual-Encoder Models for Dialogue Response Suggestions
- Learning to Represent and Predict Sets with Deep Neural Networks