Dual Attention Networks for Multimodal Reasoning and Matching
arXiv:1611.00471
Abstract
We propose Dual Attention Networks (DANs) which jointly leverage visual and textual attention mechanisms to capture fine-grained interplay between vision and language. DANs attend to specific regions in images and words in text through multiple steps and gather essential information from both modalities. Based on this framework, we introduce two types of DANs for multimodal reasoning and matching, respectively. The reasoning model allows visual and textual attentions to steer each other during collaborative inference, which is useful for tasks such as Visual Question Answering (VQA). In addition, the matching model exploits the two attention mechanisms to estimate the similarity between images and sentences by focusing on their shared semantics. Our extensive experiments validate the effectiveness of DANs in combining vision and language, achieving the state-of-the-art performance on public benchmarks for VQA and image-text matching.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- DRAW: A Recurrent Neural Network For Image Generation
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- A Hierarchical Neural Autoencoder for Paragraphs and Documents
- Show and Tell: A Neural Image Caption Generator
- Deep Networks with Internal Selective Attention through Feedback Connections
Cited by in corpus (16)
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Deep Modular Co-Attention Networks for Visual Question Answering
- High-Order Attention Models for Visual Question Answering
- Identity-Aware Textual-Visual Matching with Latent Co-attention
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- Structured Attentions for Visual Question Answering
- Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- Image-Question-Answer Synergistic Network for Visual Dialog
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- Survey of Recent Advances in Visual Question Answering
- A Simple Loss Function for Improving the Convergence and Accuracy of Visual Question Answering Models
- Automatic Music Highlight Extraction using Convolutional Recurrent Attention Networks
- Visual Explanations from Hadamard Product in Multimodal Deep Networks
- Improved RAMEN: Towards Domain Generalization for Visual Question Answering
- Applying recent advances in Visual Question Answering to Record Linkage