Stack-Captioning: Coarse-to-Fine Learning for Image Captioning
arXiv:1709.03376
Abstract
The existing image captioning approaches typically train a one-stage sentence decoder, which is difficult to generate rich fine-grained descriptions. On the other hand, multi-stage image caption model is hard to train due to the vanishing gradient problem. In this paper, we propose a coarse-to-fine multi-stage prediction framework for image captioning, composed of multiple decoders each of which operates on the output of the previous stage, producing increasingly refined image descriptions. Our proposed learning approach addresses the difficulty of vanishing gradients during training by providing a learning objective function that enforces intermediate supervisions. Particularly, we optimize our model with a reinforcement learning approach which utilizes the output of each intermediate decoder's test-time inference algorithm as well as the output of its preceding decoder to normalize the rewards, which simultaneously solves the well-known exposure bias problem and the loss-evaluation mismatch problem. We extensively evaluate the proposed approach on MSCOCO and show that our approach can achieve the state-of-the-art performance.
AAAI-2018, Oral Presentation
Cited by in corpus (23)
- A Survey of Deep Learning-based Object Detection
- Show, Recall, and Tell: Image Captioning with Recall Mechanism
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Scene Graph Generation with External Knowledge and Image Reconstruction
- Unpaired Image Captioning via Scene Graph Alignments
- Learning to Collocate Neural Modules for Image Captioning
- Fine-Grained Image Captioning with Global-Local Discriminative Objective
- Creating Something from Nothing: Unsupervised Knowledge Distillation for Cross-Modal Hashing
- Streamlined Dense Video Captioning
- Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning
- Context and Attribute Grounded Dense Captioning
- Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Watch It Twice: Video Captioning with a Refocused Video Encoder
- Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
- Geometry-Entangled Visual Semantic Transformer for Image Captioning
- Intrinsic Image Captioning Evaluation
- Zero-Shot Scene Graph Relation Prediction through Commonsense Knowledge Integration
- Hidden State Guidance: Improving Image Captioning using An Image Conditioned Autoencoder
- Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
- MUTATT: Visual-Textual Mutual Guidance for Referring Expression Comprehension
- Case Relation Transformer: A Crossmodal Language Generation Model for Fetching Instructions
- User-Aware Folk Popularity Rank: User-Popularity-Based Tag Recommendation That Can Enhance Social Popularity