Deep Reinforcement Learning-based Image Captioning with Embedding Reward
arXiv:1704.03899
Abstract
Image captioning is a challenging problem owing to the complexity in understanding the image content and diverse ways of describing it in natural language. Recent advances in deep neural networks have substantially improved the performance of this task. Most state-of-the-art approaches follow an encoder-decoder framework, which generates captions using a sequential recurrent prediction model. However, in this paper, we introduce a novel decision-making framework for image captioning. We utilize a "policy network" and a "value network" to collaboratively generate captions. The policy network serves as a local guidance by providing the confidence of predicting the next word according to the current state. Additionally, the value network serves as a global and lookahead guidance by evaluating all possible extensions of the current state. In essence, it adjusts the goal of predicting the correct words towards the goal of generating captions similar to the ground truth captions. We train both networks using an actor-critic reinforcement learning model, with a novel reward defined by visual-semantic embedding. Extensive experiments and analyses on the Microsoft COCO dataset show that the proposed framework outperforms state-of-the-art approaches across different evaluation metrics.
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Show and Tell: A Neural Image Caption Generator
- Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- CIDEr: Consensus-based Image Description Evaluation
- Simple Image Description Generator via a Linear Phrase-Based Approach
Cited by in corpus (26)
- AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
- Context-Aware Visual Policy Network for Sequence-Level Image Captioning
- Discriminability objective for training descriptive captions
- Revisiting the Master-Slave Architecture in Multi-Agent Deep Reinforcement Learning
- Reconstruction Network for Video Captioning
- Crafting a Toolchain for Image Restoration by Deep Reinforcement Learning
- No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling
- Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- A Comprehensive Survey of Deep Learning for Image Captioning
- Discourse-Aware Neural Rewards for Coherent Text Generation
- Streamlined Dense Video Captioning
- Context and Attribute Grounded Dense Captioning
- End-to-End Video Captioning with Multitask Reinforcement Learning
- Towards Amortized Ranking-Critical Training for Collaborative Filtering
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Hierarchical Photo-Scene Encoder for Album Storytelling
- Improving Reinforcement Learning Based Image Captioning with Natural Language Prior
- Tell-the-difference: Fine-grained Visual Descriptor via a Discriminating Referee
- Transfer Reward Learning for Policy Gradient-Based Text Generation
- A2-RL: Aesthetics Aware Reinforcement Learning for Image Cropping
- Improving Automatic Source Code Summarization via Deep Reinforcement Learning
- Image Captioning using Multiple Transformers for Self-Attention Mechanism
- Improving Adversarial Text Generation by Modeling the Distant Future
- FFNet: Video Fast-Forwarding via Reinforcement Learning
- Deep Reinforcement Learning with Label Embedding Reward for Supervised Image Hashing