SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
arXiv:1611.05594
Abstract
Visual attention has been successfully applied in structural prediction tasks such as visual captioning and question answering. Existing visual attention models are generally spatial, i.e., the attention is modeled as spatial probabilities that re-weight the last conv-layer feature map of a CNN encoding an input image. However, we argue that such spatial attention does not necessarily conform to the attention mechanism --- a dynamic feature extractor that combines contextual fixations over time, as CNN features are naturally spatial, channel-wise and multi-layer. In this paper, we introduce a novel convolutional neural network dubbed SCA-CNN that incorporates Spatial and Channel-wise Attentions in a CNN. In the task of image captioning, SCA-CNN dynamically modulates the sentence generation context in multi-layer feature maps, encoding where (i.e., attentive spatial locations at multiple layers) and what (i.e., attentive channels) the visual attention is. We evaluate the proposed SCA-CNN architecture on three benchmark image captioning datasets: Flickr8K, Flickr30K, and MSCOCO. It is consistently observed that SCA-CNN significantly outperforms state-of-the-art visual attention-based image captioning methods.
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- ADADELTA: An Adaptive Learning Rate Method
- Deep Residual Learning for Image Recognition
- Recurrent Models of Visual Attention
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Rethinking the Inception Architecture for Computer Vision
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Image Captioning with Semantic Attention
- Show and Tell: A Neural Image Caption Generator
- Supervised Discrete Hashing
- Deep Networks with Internal Selective Attention through Feedback Connections
- Visual Translation Embedding Network for Visual Relation Detection
- CIDEr: Consensus-based Image Description Evaluation
- Progressive Attention Networks for Visual Attribute Prediction
Cited by in corpus (19)
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Learning a Discriminative Feature Network for Semantic Segmentation
- DFANet: Deep Feature Aggregation for Real-Time Semantic Segmentation
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Relation-Aware Global Attention for Person Re-identification
- Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
- Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
- Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- Setting an attention region for convolutional neural networks using region selective features, for recognition of materials within glass vessels
- End-to-End Video Captioning with Multitask Reinforcement Learning
- Hierarchical semantic segmentation using modular convolutional neural networks
- Attention Mechanisms for Object Recognition with Event-Based Cameras
- Looking for change? Roll the Dice and demand Attention
- Deep Discrete Hashing with Self-supervised Pairwise Labels
- Meta R-CNN : Towards General Solver for Instance-level Few-shot Learning
- Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
- Report: Dynamic Eye Movement Matching and Visualization Tool in Neuro Gesture