Realistic Image Synthesis with Configurable 3D Scene Layouts
arXiv:2108.10031
Abstract
Recent conditional image synthesis approaches provide high-quality synthesized images. However, it is still challenging to accurately adjust image contents such as the positions and orientations of objects, and synthesized images often have geometrically invalid contents. To provide users with rich controllability on synthesized images in the aspect of 3D geometry, we propose a novel approach to realistic-looking image synthesis based on a configurable 3D scene layout. Our approach takes a 3D scene with semantic class labels as input and trains a 3D scene painting network that synthesizes color values for the input 3D scene. With the trained painting network, realistic-looking images for the input 3D scene can be rendered and manipulated. To train the painting network without 3D color supervision, we exploit an off-the-shelf 2D semantic image synthesis method. In experiments, we show that our approach produces images with geometrically correct structures and supports geometric manipulation such as the change of the viewpoint and object poses as well as manipulation of the painting style.
paper: 9 pages, supplementary materials: 7 pages
References in corpus (5)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
- Zero-Shot Text-to-Image Generation
- Learning to Predict Layout-to-image Conditional Convolutions for Semantic Image Synthesis
- You Only Need Adversarial Supervision for Semantic Image Synthesis