Attribute2Image: Conditional Image Generation from Visual Attributes
arXiv:1512.00570
Abstract
This paper investigates a novel problem of generating images from visual attributes. We model the image as a composite of foreground and background and develop a layered generative model with disentangled latent variables that can be learned end-to-end using a variational auto-encoder. We experiment with natural images of faces and birds and demonstrate that the proposed models are capable of generating realistic and diverse samples with disentangled latent representations. We use a general energy minimization algorithm for posterior inference of latent variables given novel images. Therefore, the learned generative models show excellent quantitative and visual results in the tasks of attribute-conditioned image reconstruction and completion.
19 pages, accepted by ECCV 2016, The 14th European Conference on Computer Vision (2016)
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Understanding Neural Networks Through Deep Visualization
- DRAW: A Recurrent Neural Network For Image Generation
- Deep Convolutional Inverse Graphics Network
- Text to 3D Scene Generation with Rich Lexical Grounding
- Transferring Landmark Annotations for Cross-Dataset Face Alignment
Cited by in corpus (41)
- Learning Discourse-level Diversity for Neural Dialog Models using Conditional Variational Autoencoders
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Semantic Autoencoder for Zero-Shot Learning
- Person Transfer GAN to Bridge Domain Gap for Person Re-Identification
- Coupled Generative Adversarial Networks
- Photographic Image Synthesis with Cascaded Refinement Networks
- Transformation-Based Models of Video Sequences
- CVAE-GAN: Fine-Grained Image Generation through Asymmetric Training
- Convolutional Network for Attribute-driven and Identity-preserving Human Face Generation
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Stacked Generative Adversarial Networks
- Tackling Over-pruning in Variational Autoencoders
- Multi-View Image Generation from a Single-View
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Conditional Generative Moment-Matching Networks
- Towards Open-Set Identity Preserving Face Synthesis
- Creativity: Generating Diverse Questions using Variational Autoencoders
- A Multi-Modal Chinese Poetry Generation Model
- Pose-Normalized Image Generation for Person Re-identification
- Scribbler: Controlling Deep Image Synthesis with Sketch and Color
- Face Attribute Prediction Using Off-the-Shelf CNN Features
- Precomputed Real-Time Texture Synthesis with Markovian Generative Adversarial Networks
- Localizing by Describing: Attribute-Guided Attention Localization for Fine-Grained Recognition
- Predicting Visual Exemplars of Unseen Classes for Zero-Shot Learning
- Ranking CGANs: Subjective Control over Semantic Image Attributes
- Facelet-Bank for Fast Portrait Manipulation
- Deep Feature Consistent Variational Autoencoder
- DA-GAN: Instance-level Image Translation by Deep Attention Generative Adversarial Networks (with Supplementary Materials)
- Bottleneck Conditional Density Estimation
- Multi-Domain Level Generation and Blending with Sketches via Example-Driven BSP and Variational Autoencoders
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Smart, Sparse Contours to Represent and Edit Images
- Class label autoencoder for zero-shot learning
- Improving Bi-directional Generation between Different Modalities with Variational Autoencoders
- Depth Structure Preserving Scene Image Generation
- Conditional Level Generation and Game Blending
- DISCO Nets: DISsimilarity COefficient Networks
- Focus-Constrained Attention Mechanism for CVAE-based Response Generation
- VConstruct: Filling Gaps in Chl-a Data Using a Variational Autoencoder
- LR-to-HR Face Hallucination with an Adversarial Progressive Attribute-Induced Network
- Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold