Learning What and Where to Draw
arXiv:1610.02454
Abstract
Generative Adversarial Networks (GANs) have recently demonstrated the capability to synthesize compelling real-world images, such as room interiors, album covers, manga, faces, birds, and flowers. While existing models can synthesize images based on global constraints such as a class label or caption, they do not provide control over pose or object location. We propose a new model, the Generative Adversarial What-Where Network (GAWWN), that synthesizes images given instructions describing what content to draw in which location. We show high-quality 128 x 128 image synthesis on the Caltech-UCSD Birds dataset, conditioned on both informal text descriptions and also object location. Our system exposes control over both the bounding box around the bird and its constituent parts. By modeling the conditional distributions over part locations, our system also enables conditioning on arbitrary subsets of parts (e.g. only the beak and tail), yielding an efficient interface for picking part locations. We also show preliminary results on the more challenging domain of text- and location-controllable synthesis of images of human actions on the MPII Human Pose dataset.
In NIPS 2016
References in corpus (5)
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- DRAW: A Recurrent Neural Network For Image Generation
- Deep Convolutional Inverse Graphics Network
Cited by in corpus (5)
- Spatial Evolutionary Generative Adversarial Networks
- Parallel Multiscale Autoregressive Density Estimation
- Improving Consistency and Correctness of Sequence Inpainting using Semantically Guided Generative Adversarial Network
- Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching
- ReshapeGAN: Object Reshaping by Providing A Single Reference Image