Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
arXiv:1612.01887
Abstract
Attention-based neural encoder-decoder frameworks have been widely adopted for image captioning. Most methods force visual attention to be active for every generated word. However, the decoder likely requires little to no visual information from the image to predict non-visual words such as "the" and "of". Other words that may seem visual can often be predicted reliably just from the language model e.g., "sign" after "behind a red stop" or "phone" following "talking on a cell". In this paper, we propose a novel adaptive attention model with a visual sentinel. At each time step, our model decides whether to attend to the image (and if so, to which regions) or to the visual sentinel. The model decides whether to attend to the image and where, in order to extract meaningful information for sequential word generation. We test our method on the COCO image captioning 2015 challenge dataset and Flickr30K. Our approach sets the new state-of-the-art by a significant margin.
12 pages, 11 figures, CVPR2017 camera ready
References in corpus (4)
Cited by in corpus (26)
- Bayesian Recurrent Neural Networks
- Towards Diverse and Natural Image Descriptions via a Conditional GAN
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Discriminability objective for training descriptive captions
- Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
- Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
- A Neural Compositional Paradigm for Image Captioning
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Fast Image Caption Generation with Position Alignment
- MemexQA: Visual Memex Question Answering
- Sketch-R2CNN: An Attentive Network for Vector Sketch Recognition
- An Attempt towards Interpretable Audio-Visual Video Captioning
- Dynamic Spatial-Temporal Representation Learning for Traffic Flow Prediction
- Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Sparse Attentive Backtracking: Long-Range Credit Assignment in Recurrent Networks
- Two-Stage Synthesis Networks for Transfer Learning in Machine Comprehension
- Netizen-Style Commenting on Fashion Photos: Dataset and Diversity Measures
- Improving Reinforcement Learning Based Image Captioning with Natural Language Prior
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Towards Diverse Paragraph Captioning for Untrimmed Videos
- Generating Descriptions for Sequential Images with Local-Object Attention and Global Semantic Context Modelling
- Morphological Skip-Gram: Using morphological knowledge to improve word representation
- Better Understanding Hierarchical Visual Relationship for Image Caption