ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering
arXiv:1511.05960
Abstract
We propose a novel attention based deep learning architecture for visual question answering task (VQA). Given an image and an image related natural language question, VQA generates the natural language answer for the question. Generating the correct answers requires the model's attention to focus on the regions corresponding to the question, because different questions inquire about the attributes of different image regions. We introduce an attention based configurable convolutional neural network (ABC-CNN) to learn such question-guided attention. ABC-CNN determines an attention map for an image-question pair by convolving the image feature map with configurable convolutional kernels derived from the question's semantics. We evaluate the ABC-CNN architecture on three benchmark VQA datasets: Toronto COCO-QA, DAQUAR, and VQA dataset. ABC-CNN model achieves significant improvements over state-of-the-art methods on these datasets. The question-guided attention generated by ABC-CNN is also shown to reflect the regions that are highly relevant to the questions.
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- VQA: Visual Question Answering
- Recurrent Models of Visual Attention
- Fully Convolutional Networks for Semantic Segmentation
- Multiple Object Recognition with Visual Attention
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
- Attention for Fine-Grained Categorization
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Bilinear CNNs for Fine-grained Visual Recognition
- Automatic Concept Discovery from Parallel Text and Visual Corpora
Cited by in corpus (65)
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering
- This Looks Like That: Deep Learning for Interpretable Image Recognition
- Simple Baseline for Visual Question Answering
- Temporal Attention augmented Bilinear Network for Financial Time-Series Data Analysis
- TI-CNN: Convolutional Neural Networks for Fake News Detection
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- Reinforcement Learning for Solving the Vehicle Routing Problem
- Attention to Scale: Scale-aware Semantic Image Segmentation
- A Focused Dynamic Attention Model for Visual Question Answering
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Deep Modular Co-Attention Networks for Visual Question Answering
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- From Easy to Hard: Learning Language-guided Curriculum for Visual Question Answering on Remote Sensing Data
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Attention Correctness in Neural Image Captioning
- Attentive Explanations: Justifying Decisions and Pointing to the Evidence
- ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence Pairs
- Vision-to-Language Tasks Based on Attributes and Attention Mechanism
- Visual Question Answering: A Survey of Methods and Datasets
- Implicit Distortion and Fertility Models for Attention-based Encoder-Decoder NMT Model
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- FVQA: Fact-based Visual Question Answering
- Multimodal Unified Attention Networks for Vision-and-Language Interactions
- Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Lexicon Integrated CNN Models with Attention for Sentiment Analysis
- Question-Guided Hybrid Convolution for Visual Question Answering
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- Dynamic Computational Time for Visual Attention
- How to find a good image-text embedding for remote sensing visual question answering?
- VQABQ: Visual Question Answering by Basic Questions
- Robustness Analysis of Visual QA Models by Basic Questions
- AMC: Attention guided Multi-modal Correlation Learning for Image Search
- Understanding the computational demands underlying visual reasoning
- Setting an attention region for convolutional neural networks using region selective features, for recognition of materials within glass vessels
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- Aspect-augmented Adversarial Networks for Domain Adaptation
- Deep Graph Attention Model
- Mining Significant Microblogs for Misinformation Identification: An Attention-based Approach
- An Empirical Evaluation of Visual Question Answering for Novel Objects
- MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering
- Learning Visual Question Answering by Bootstrapping Hard Attention
- Compact Global Descriptor for Neural Networks
- Large Scale Multimodal Classification Using an Ensemble of Transformer Models and Co-Attention
- MarioQA: Answering Questions by Watching Gameplay Videos
- Classifying a specific image region using convolutional nets with an ROI mask as input
- Zero-Shot Transfer VQA Dataset
- Dual Recurrent Attention Units for Visual Question Answering
- LS-Tree: Model Interpretation When the Data Are Linguistic
- Learning Semantically Coherent and Reusable Kernels in Convolution Neural Nets for Sentence Classification
- Transformers for Limit Order Books
- Hierarchical semantic segmentation using modular convolutional neural networks
- MAANet: Multi-view Aware Attention Networks for Image Super-Resolution
- Attention-based Memory Selection Recurrent Network for Language Modeling
- Exploring Different Dimensions of Attention for Uncertainty Detection
- TAB-VCR: Tags and Attributes based Visual Commonsense Reasoning Baselines
- Locally Smoothed Neural Networks
- Learning Sparse Mixture of Experts for Visual Question Answering
- Representation Learning for Grounded Spatial Reasoning
- Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization
- Improved RAMEN: Towards Domain Generalization for Visual Question Answering
- Transfer Learning in Visual and Relational Reasoning
- Recognizing Part Attributes with Insufficient Data
- TexRel: a Green Family of Datasets for Emergent Communications on Relations