Attention Correctness in Neural Image Captioning
arXiv:1605.09553
Abstract
Attention mechanisms have recently been introduced in deep learning for various tasks in natural language processing and computer vision. But despite their popularity, the "correctness" of the implicitly-learned attention maps has only been assessed qualitatively by visualization of several examples. In this paper we focus on evaluating and improving the correctness of attention in neural image captioning models. Specifically, we propose a quantitative evaluation metric for the consistency between the generated attention maps and human annotations, using recently released datasets with alignment between regions in images and entities in captions. We then propose novel models with different levels of explicit supervision for learning attention maps during training. The supervision can be strong when alignment between regions and caption entities are available, or weak when only object segments and categories are provided. We show on the popular Flickr30k and COCO datasets that introducing supervision of attention maps during training solidly improves both attention correctness and caption quality, showing the promise of making machine perception more human-like.
To appear in AAAI-17. See http://www.cs.jhu.edu/~cxliu/ for supplementary material
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Image Captioning with Semantic Attention
- Learning a Recurrent Visual Representation for Image Caption Generation
Cited by in corpus (28)
- Understanding Attention and Generalization in Graph Neural Networks
- Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
- Counterfactual Samples Synthesizing for Robust Visual Question Answering
- TVQA+: Spatio-Temporal Grounding for Video Question Answering
- Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
- Semantic Compositional Networks for Visual Captioning
- Revisiting Image-Language Networks for Open-ended Phrase Detection
- Generating Descriptions with Grounded and Co-Referenced People
- Learning to Globally Edit Images with Textual Description
- Adversarial TableQA: Attention Supervision for Question Answering on Tables
- Neural Sign Language Translation based on Human Keypoint Estimation
- MAT: A Multimodal Attentive Translator for Image Captioning
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- Scene Graph Parsing by Attention Graph
- Scene Graph Parsing as Dependency Parsing
- Multi-level Multimodal Common Semantic Space for Image-Phrase Grounding
- Top-down Visual Saliency Guided by Captions
- An Empirical Study of Language CNN for Image Captioning
- Teaching Machines to Code: Neural Markup Generation with Visual Attention
- Mitigating Gender Bias in Captioning Systems
- Grounded Video Description
- Transfer Reward Learning for Policy Gradient-Based Text Generation
- Aligned Image-Word Representations Improve Inductive Transfer Across Vision-Language Tasks
- Recurrent Multimodal Interaction for Referring Image Segmentation
- Semantic Grouping Network for Video Captioning
- Image Captioning with Integrated Bottom-Up and Multi-level Residual Top-Down Attention for Game Scene Understanding
- Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real Images