Diversity-Sensitive Conditional Generative Adversarial Networks
arXiv:1901.09024
Abstract
We propose a simple yet highly effective method that addresses the mode-collapse problem in the Conditional Generative Adversarial Network (cGAN). Although conditional distributions are multi-modal (i.e., having many modes) in practice, most cGAN approaches tend to learn an overly simplified distribution where an input is always mapped to a single output regardless of variations in latent code. To address such issue, we propose to explicitly regularize the generator to produce diverse outputs depending on latent codes. The proposed regularization is simple, general, and can be easily integrated into most conditional GAN objectives. Additionally, explicit regularization on generator allows our method to control a balance between visual quality and diversity. We demonstrate the effectiveness of our method on three conditional generation tasks: image-to-image translation, image inpainting, and future video prediction. We show that simple addition of our regularization to existing models leads to surprisingly diverse generations, substantially outperforming the previous approaches for multi-modal conditional generation specifically designed in each individual task.
Accepted as a conference paper at ICLR 2019
Cited by in corpus (14)
- Physics-Constrained Deep Learning for High-dimensional Surrogate Modeling and Uncertainty Quantification without Labeled Data
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- Few-shot Image Generation via Cross-domain Correspondence
- SSD-GAN: Measuring the Realness in the Spatial and Spectral Domains
- Regularizing Generative Adversarial Networks under Limited Data
- Diverse Semantic Image Synthesis via Probability Distribution Modeling
- Max-Affine Spline Insights into Deep Generative Networks
- Image-to-Image Translation with Low Resolution Conditioning
- Conditional Coupled Generative Adversarial Networks for Zero-Shot Domain Adaptation
- A Generative Model for Hallucinating Diverse Versions of Super Resolution Images
- SLGAN: Style- and Latent-guided Generative Adversarial Network for Desirable Makeup Transfer and Removal
- Improving Text to Image Generation using Mode-seeking Function
- Multimodal Image Synthesis with Conditional Implicit Maximum Likelihood Estimation
- Toward Zero-Shot Unsupervised Image-to-Image Translation