Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
arXiv:1703.10476
Abstract
While strong progress has been made in image captioning over the last years, machine and human captions are still quite distinct. A closer look reveals that this is due to the deficiencies in the generated word distribution, vocabulary size, and strong bias in the generators towards frequent captions. Furthermore, humans -- rightfully so -- generate multiple, diverse captions, due to the inherent ambiguity in the captioning task which is not considered in today's systems. To address these challenges, we change the training objective of the caption generator from reproducing groundtruth captions to generating a set of captions that is indistinguishable from human generated captions. Instead of handcrafting such a learning target, we employ adversarial training in combination with an approximate Gumbel sampler to implicitly match the generated distribution to the human one. While our method achieves comparable performance to the state-of-the-art in terms of the correctness of the captions, we generate a set of diverse captions, that are significantly less biased and match the word statistics better in several aspects.
16 pages, Published in ICCV 2017
References in corpus (29)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Deep Residual Learning for Image Recognition
- Categorical Reparameterization with Gumbel-Softmax
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Improved Techniques for Training GANs
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Sequence Level Training with Recurrent Neural Networks
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Semantic Segmentation using Adversarial Networks
- Improved Image Captioning via Policy Gradient optimization of SPIDEr
- Deep multi-scale video prediction beyond mean square error
- SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
- A Diversity-Promoting Objective Function for Neural Conversation Models
- Image Captioning with Semantic Attention
- Adversarial Learning for Neural Dialogue Generation
- How (not) to Train your Generative Model: Scheduled Sampling, Likelihood, Adversary?
- Show and Tell: A Neural Image Caption Generator
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Language Models for Image Captioning: The Quirks and What Works
- Towards Diverse and Natural Image Descriptions via a Conditional GAN
- Generating Visual Explanations
- CIDEr: Consensus-based Image Description Evaluation
- Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
- Boosting Image Captioning with Attributes
- Reasoning About Pragmatics with Neural Listeners and Speakers
- Creativity: Generating Diverse Questions using Variational Autoencoders
Cited by in corpus (17)
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Language Generation with Recurrent Generative Adversarial Networks without Pre-training
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Discriminability objective for training descriptive captions
- Equal But Not The Same: Understanding the Implicit Relationship Between Persuasive Images and Text
- pix2code: Generating Code from a Graphical User Interface Screenshot
- Describing like humans: on diversity in image captioning
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Learning to Evaluate Image Captioning
- A Comprehensive Survey of Deep Learning for Image Captioning
- Paying More Attention to Saliency: Image Captioning with Saliency and Context Attention
- Improving Conditional Sequence Generative Adversarial Networks by Stepwise Evaluation
- Latent Normalizing Flows for Many-to-Many Cross-Domain Mappings
- : Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
- Improving Reinforcement Learning Based Image Captioning with Natural Language Prior
- Group-based Distinctive Image Captioning with Memory Attention
- Towards Diverse Paragraph Captioning for Untrimmed Videos