High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs
arXiv:1711.11585
Abstract
We present a new method for synthesizing high-resolution photo-realistic images from semantic label maps using conditional generative adversarial networks (conditional GANs). Conditional GANs have enabled a variety of applications, but the results are often limited to low-resolution and still far from realistic. In this work, we generate 2048x1024 visually appealing results with a novel adversarial loss, as well as new multi-scale generator and discriminator architectures. Furthermore, we extend our framework to interactive visual manipulation with two additional features. First, we incorporate object instance segmentation information, which enables object manipulations such as removing/adding objects and changing the object category. Second, we propose a method to generate diverse results given the same input, allowing users to edit the object appearance interactively. Human opinion studies demonstrate that our method significantly outperforms existing methods, advancing both the quality and the resolution of deep image synthesis and editing.
v2: CVPR camera ready, adding more results for edge-to-photo examples
Cited by in corpus (37)
- Progressive Growing of GANs for Improved Quality, Stability, and Variation
- A survey on GANs for computer vision: Recent research, analysis and taxonomy
- I2V-GAN: Unpaired Infrared-to-Visible Video Translation
- Generative Modeling of Turbulence
- Patch-based Progressive 3D Point Set Upsampling
- Resolution enhancement and realistic speckle recovery with generative adversarial modeling of micro-optical coherence tomography
- DAMO: Deep Agile Mask Optimization for Full Chip Scale
- Text as Neural Operator: Image Manipulation by Text Instruction
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
- Fast and Scalable Earth Texture Synthesis using Spatially Assembled Generative Adversarial Neural Networks
- Human Motion Transfer with 3D Constraints and Detail Enhancement
- Automatic vocal tract landmark localization from midsagittal MRI data
- PanoGAN A Deep Generative Model for Panoramic Dental Radiographs
- Local Class-Specific and Global Image-Level Generative Adversarial Networks for Semantic-Guided Scene Generation
- UGAN: Untraceable GAN for Multi-Domain Face Translation
- Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
- Disentangling Propagation and Generation for Video Prediction
- COCO-GAN: Generation by Parts via Conditional Coordinating
- Image Generation from Layout
- Robust Pose Transfer with Dynamic Details using Neural Video Rendering
- StarSRGAN: Improving Real-World Blind Super-Resolution
- CoNeS: Conditional neural fields with shift modulation for multi-sequence MRI translation
- Fine Structure-Aware Sampling: A New Sampling Training Scheme for Pixel-Aligned Implicit Models in Single-View Human Reconstruction
- Deep Network Interpolation for Continuous Imagery Effect Transition
- CompressNet: Generative Compression at Extremely Low Bitrates
- FaceShapeGene: A Disentangled Shape Representation for Flexible Face Image Editing
- Im2Pencil: Controllable Pencil Illustration from Photographs
- Enforcing Perceptual Consistency on Generative Adversarial Networks by Using the Normalised Laplacian Pyramid Distance
- Sparsely Grouped Multi-task Generative Adversarial Networks for Facial Attribute Manipulation
- How Old Are You? Face Age Translation with Identity Preservation Using GANs
- Modeling Deep Learning Based Privacy Attacks on Physical Mail
- Open-World Entity Segmentation
- Video-to-Video Translation for Visual Speech Synthesis
- Synthesizing Photorealistic Images with Deep Generative Learning
- A self-adapting super-resolution structures framework for automatic design of GAN
- Linear Semantics in Generative Adversarial Networks
- ReshapeGAN: Object Reshaping by Providing A Single Reference Image