Context-Aware Visual Policy Network for Sequence-Level Image Captioning
arXiv:1808.05864 · doi:10.1145/3240508.3240632
Abstract
Many vision-language tasks can be reduced to the problem of sequence prediction for natural language output. In particular, recent advances in image captioning use deep reinforcement learning (RL) to alleviate the "exposure bias" during training: ground-truth subsequence is exposed in every step prediction, which introduces bias in test when only predicted subsequence is seen. However, existing RL-based image captioning methods only focus on the language policy while not the visual policy (e.g., visual attention), and thus fail to capture the visual context that are crucial for compositional reasoning such as visual relationships (e.g., "man riding horse") and comparisons (e.g., "smaller cat"). To fill the gap, we propose a Context-Aware Visual Policy network (CAVP) for sequence-level image captioning. At every time step, CAVP explicitly accounts for the previous visual attentions as the context, and then decides whether the context is helpful for the current word generation given the current visual attention. Compared against traditional visual attention that only fixes a single image region at every step, CAVP can attend to complex visual compositions over time. The whole image captioning model --- CAVP and its subsequent language policy network --- can be efficiently optimized end-to-end by using an actor-critic policy gradient method with respect to any caption evaluation metric. We demonstrate the effectiveness of CAVP by state-of-the-art performances on MS-COCO offline split and online server, using various metrics and sensible visualizations of qualitative visual context. The code is available at https://github.com/daqingliu/CAVP
9 pages, 6 figures, ACM MM 2018 oral
References in corpus (4)
Cited by in corpus (23)
- Deconfounded Image Captioning: A Causal Retrospect
- Image Captioning: Transforming Objects into Words
- Show, Recall, and Tell: Image Captioning with Recall Mechanism
- Improving Image Captioning with Better Use of Captions
- A Better Variant of Self-Critical Sequence Training
- Normalized and Geometry-Aware Self-Attention Network for Image Captioning
- Auto-Encoding Scene Graphs for Image Captioning
- Counterfactual Critic Multi-Agent Training for Scene Graph Generation
- Learning to Collocate Neural Modules for Image Captioning
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- Curiosity-driven Reinforcement Learning for Diverse Visual Paragraph Generation
- Learning to Discretely Compose Reasoning Module Networks for Video Captioning
- Causal Attention for Vision-Language Tasks
- Unpaired Cross-lingual Image Caption Generation with Self-Supervised Rewards
- CaseNet: Content-Adaptive Scale Interaction Networks for Scene Parsing
- Image Captioning with Context-Aware Auxiliary Guidance
- Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning
- New Ideas and Trends in Deep Multimodal Content Understanding: A Review
- Exploring and Distilling Cross-Modal Information for Image Captioning
- Iterative Context-Aware Graph Inference for Visual Dialog
- Goal-driven text descriptions for images
- PR Product: A Substitute for Inner Product in Neural Networks
- Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features