Plug & Play Generative Networks: Conditional Iterative Generation of Images in Latent Space
arXiv:1612.00005
Abstract
Generating high-resolution, photo-realistic images has been a long-standing goal in machine learning. Recently, Nguyen et al. (2016) showed one interesting way to synthesize novel images by performing gradient ascent in the latent space of a generator network to maximize the activations of one or multiple neurons in a separate classifier network. In this paper we extend this method by introducing an additional prior on the latent code, improving both sample quality and sample diversity, leading to a state-of-the-art generative model that produces high quality images at higher resolutions (227x227) than previous generative models, and does so for all 1000 ImageNet categories. In addition, we provide a unified probabilistic interpretation of related activation maximization methods and call the general class of models "Plug and Play Generative Networks". PPGNs are composed of 1) a generator network G that is capable of drawing a wide range of image types and 2) a replaceable "condition" network C that tells the generator what to draw. We demonstrate the generation of images conditioned on a class (when C is an ImageNet or MIT Places classification network) and also conditioned on a caption (when C is an image captioning network). Our method also improves the state of the art of Multifaceted Feature Visualization, which generates the set of synthetic inputs that activate a neuron in order to better understand how deep neural networks operate. Finally, we show that our model performs reasonably well at the task of image inpainting. While image models are used in this paper, the approach is modality-agnostic and can be applied to many types of data.
CVPR camera-ready
Cited by in corpus (32)
- NIPS 2016 Tutorial: Generative Adversarial Networks
- Unsupervised Image-to-Image Translation Networks
- How Generative Adversarial Networks and Their Variants Work: An Overview
- Lung and Pancreatic Tumor Characterization in the Deep Learning Era: Novel Supervised and Unsupervised Learning Approaches
- BAGAN: Data Augmentation with Balancing GAN
- On the Reconstruction of Face Images from Deep Face Templates
- Bayesian parameter estimation using conditional variational autoencoders for gravitational-wave astronomy
- Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks
- An Introduction to Image Synthesis with Generative Adversarial Nets
- Design of metalloproteins and novel protein folds using variational autoencoders
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- Deep Bayesian Inversion
- CVAE-GAN: Fine-Grained Image Generation through Asymmetric Training
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Distribution Matching in Variational Inference
- Activation Maximization Generative Adversarial Nets
- Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings
- Semi-Latent GAN: Learning to generate and modify facial images from attributes
- Gang of GANs: Generative Adversarial Networks with Maximum Margin Ranking
- ReenactGAN: Learning to Reenact Faces via Boundary Transfer
- On the Effectiveness of Least Squares Generative Adversarial Networks
- Robustifying Models Against Adversarial Attacks by Langevin Dynamics
- Generating Hard Examples for Pixel-wise Classification
- Understanding the Effectiveness of Lipschitz-Continuity in Generative Adversarial Nets
- Contextual-based Image Inpainting: Infer, Match, and Translate
- GridFace: Face Rectification via Learning Local Homography Transformations
- Smart, Sparse Contours to Represent and Edit Images
- Bayesian Reasoning with Trained Neural Networks
- Lipschitz Constrained GANs via Boundedness and Continuity
- Learning Discriminators as Energy Networks in Adversarial Learning
- Understanding Regularization to Visualize Convolutional Neural Networks
- Updating the generator in PPGN-h with gradients flowing through the encoder