Towards Diverse and Natural Image Descriptions via a Conditional GAN
arXiv:1703.06029
Abstract
Despite the substantial progress in recent years, the image captioning techniques are still far from being perfect.Sentences produced by existing methods, e.g. those based on RNNs, are often overly rigid and lacking in variability. This issue is related to a learning principle widely used in practice, that is, to maximize the likelihood of training samples. This principle encourages high resemblance to the "ground-truth" captions while suppressing other reasonable descriptions. Conventional evaluation metrics, e.g. BLEU and METEOR, also favor such restrictive methods. In this paper, we explore an alternative approach, with the aim to improve the naturalness and diversity -- two essential properties of human expression. Specifically, we propose a new framework based on Conditional Generative Adversarial Networks (CGAN), which jointly learns a generator to produce descriptions conditioned on images and an evaluator to assess how well a description fits the visual content. It is noteworthy that training a sequence generator is nontrivial. We overcome the difficulty by Policy Gradient, a strategy stemming from Reinforcement Learning, which allows the generator to receive early feedback along the way. We tested our method on two large datasets, where it performed competitively against real people in our user study and outperformed other methods on various tasks.
accepted in ICCV2017 as an Oral paper
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Conditional Generative Adversarial Nets
- Improved Image Captioning via Policy Gradient optimization of SPIDEr
- Show and Tell: A Neural Image Caption Generator
- Exploring Nearest Neighbor Approaches for Image Captioning
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Detecting Visual Relationships with Deep Relational Networks
- CIDEr: Consensus-based Image Description Evaluation
- Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
- Simple Image Description Generator via a Linear Phrase-Based Approach
Cited by in corpus (41)
- Generative Adversarial Network in Medical Imaging: A Review
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Improving Image Captioning with Conditional Generative Adversarial Nets
- Discriminability objective for training descriptive captions
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- EvalAI: Towards Better Evaluation Systems for AI Agents
- A Neural Compositional Paradigm for Image Captioning
- Teaching Machines to Describe Images via Natural Language Feedback
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- pix2code: Generating Code from a Graphical User Interface Screenshot
- Scene Graph Generation from Objects, Phrases and Region Captions
- Masked Non-Autoregressive Image Captioning
- Text-to-Image-to-Text Translation using Cycle Consistent Adversarial Networks
- No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling
- Describing like humans: on diversity in image captioning
- Overcoming Language Priors in Visual Question Answering with Adversarial Regularization
- Learning to Globally Edit Images with Textual Description
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Learning to Evaluate Image Captioning
- Latent Translation: Crossing Modalities by Bridging Generative Models
- A Comprehensive Survey of Deep Learning for Image Captioning
- Learning to Act Properly: Predicting and Explaining Affordances from Images
- Improving Movement Predictions of Traffic Actors in Bird's-Eye View Models using GANs and Differentiable Trajectory Rasterization
- Move Forward and Tell: A Progressive Generator of Video Descriptions
- End-to-End Video Captioning with Multitask Reinforcement Learning
- Latent Normalizing Flows for Many-to-Many Cross-Domain Mappings
- Macroscopic Control of Text Generation for Image Captioning
- Neural Text Generation with Artificial Negative Examples
- Blind Motion Deblurring with Cycle Generative Adversarial Networks
- Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation
- Improving Reinforcement Learning Based Image Captioning with Natural Language Prior
- Unsupervised Stylish Image Description Generation via Domain Layer Norm
- Learning to Caption Images through a Lifetime by Asking Questions
- EnsembleGAN: Adversarial Learning for Retrieval-Generation Ensemble Model on Short-Text Conversation
- Self-Annotated Training for Controllable Image Captioning
- A Face-to-Face Neural Conversation Model
- Group-based Distinctive Image Captioning with Memory Attention
- Escaping from Collapsing Modes in a Constrained Space
- Middle-Out Decoding