Training Recurrent Answering Units with Joint Loss Minimization for VQA
arXiv:1606.03647
Abstract
We propose a novel algorithm for visual question answering based on a recurrent deep neural network, where every module in the network corresponds to a complete answering unit with attention mechanism by itself. The network is optimized by minimizing loss aggregated from all the units, which share model parameters while receiving different information to compute attention probability. For training, our model attends to a region within image feature map, updates its memory based on the question and attended image feature, and answers the question based on its memory state. This procedure is performed to compute loss in each step. The motivation of this approach is our observation that multi-step inferences are often required to answer questions while each problem may have a unique desirable number of steps, which is difficult to identify in practice. Hence, we always make the first unit in the network solve problems, but allow it to learn the knowledge from the rest of units by backpropagation unless it degrades the model. To implement this idea, we early-stop training each unit as soon as it starts to overfit. Note that, since more complex models tend to overfit on easier questions quickly, the last answering unit in the unfolded recurrent neural network is typically killed first while the first one remains last. We make a single-step prediction for a new question using the shared model. This strategy works better than the other options within our framework since the selected model is trained effectively from all units without overfitting. The proposed algorithm outperforms other multi-step attention based approaches using a single step prediction in VQA dataset.
References in corpus (4)
Cited by in corpus (23)
- Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering
- Visual Question Answering: Datasets, Algorithms, and Future Challenges
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Visual Reference Resolution using Attention Memory for Visual Dialog
- C-VQA: A Compositional Split of the Visual Question Answering (VQA) v1.0 Dataset
- Dual Attention Networks for Multimodal Reasoning and Matching
- High-Order Attention Models for Visual Question Answering
- Visual Question Answering: A Survey of Methods and Datasets
- Multi-modality Latent Interaction Network for Visual Question Answering
- DVQA: Understanding Data Visualizations via Question Answering
- Structured Attentions for Visual Question Answering
- MUREL: Multimodal Relational Reasoning for Visual Question Answering
- An Analysis of Visual Question Answering Algorithms
- Question-Guided Hybrid Convolution for Visual Question Answering
- Visual Question Answering with Memory-Augmented Networks
- Compact Trilinear Interaction for Visual Question Answering
- The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)
- Question Type Guided Attention in Visual Question Answering
- Dual Recurrent Attention Units for Visual Question Answering
- Answer Them All! Toward Universal Visual Question Answering Models
- Text-guided Attention Model for Image Captioning
- Chop Chop BERT: Visual Question Answering by Chopping VisualBERT's Heads
- Improved RAMEN: Towards Domain Generalization for Visual Question Answering