Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner
arXiv:1705.00930
Abstract
Impressive image captioning results are achieved in domains with plenty of training image and sentence pairs (e.g., MSCOCO). However, transferring to a target domain with significant domain shifts but no paired training data (referred to as cross-domain image captioning) remains largely unexplored. We propose a novel adversarial training procedure to leverage unpaired data in the target domain. Two critic networks are introduced to guide the captioner, namely domain critic and multi-modal critic. The domain critic assesses whether the generated sentences are indistinguishable from sentences in the target domain. The multi-modal critic assesses whether an image and its generated sentence are a valid pair. During training, the critics and captioner act as adversaries -- captioner aims to generate indistinguishable sentences, whereas critics aim at distinguishing them. The assessment improves the captioner through policy gradient updates. During inference, we further propose a novel critic-based planning method to select high-quality sentences without additional supervision (e.g., tags). To evaluate, we use MSCOCO as the source domain and four other datasets (CUB-200-2011, Oxford-102, TGIF, and Flickr30k) as the target domains. Our method consistently performs well on all datasets. In particular, on CUB-200-2011, we achieve 21.8% CIDEr-D improvement after adaptation. Utilizing critics during inference further gives another 4.5% boost.
ICCV 2017
References in corpus (13)
- Deep Residual Learning for Image Recognition
- Learning Transferable Features with Deep Adaptation Networks
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Sequence Level Training with Recurrent Neural Networks
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Domain-Adversarial Neural Networks
- An Actor-Critic Algorithm for Sequence Prediction
- Central Moment Discrepancy (CMD) for Domain-Invariant Representation Learning
- Show and Tell: A Neural Image Caption Generator
- Generating Visual Explanations
- Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data
- Title Generation for User Generated Videos
Cited by in corpus (10)
- Deep Visual Domain Adaptation: A Survey
- An Introduction to Image Synthesis with Generative Adversarial Nets
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach
- Early Action Prediction with Generative Adversarial Networks
- CariGAN: Caricature Generation through Weakly Paired Adversarial Learning
- CANZSL: Cycle-Consistent Adversarial Networks for Zero-Shot Learning from Natural Language
- A Comprehensive Survey of Deep Learning for Image Captioning
- Style Obfuscation by Invariance
- Unsupervised Stylish Image Description Generation via Domain Layer Norm