Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
arXiv:1703.02921
Abstract
We present a transformation-grounded image generation network for novel 3D view synthesis from a single image. Instead of taking a 'blank slate' approach, we first explicitly infer the parts of the geometry visible both in the input and novel views and then re-cast the remaining synthesis problem as image completion. Specifically, we both predict a flow to move the pixels from the input to the novel view along with a novel visibility map that helps deal with occulsion/disocculsion. Next, conditioned on those intermediate results, we hallucinate (infer) parts of the object invisible in the input image. In addition to the new network structure, training with a combination of adversarial and perceptual loss results in a reduction in common artifacts of novel view synthesis such as distortions and holes, while successfully generating high frequency details and preserving visual aspects of the input image. We evaluate our approach on a wide range of synthetic and real examples. Both qualitative and quantitative results show our method achieves significantly better results compared to existing methods.
To appear in CVPR 2017
References in corpus (10)
- Conditional Generative Adversarial Nets
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- Generating Videos with Scene Dynamics
- Deep Convolutional Inverse Graphics Network
- Learning Deconvolution Network for Semantic Segmentation
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Render for CNN: Viewpoint Estimation in Images Using CNNs Trained with Rendered 3D Model Views
- DeepStereo: Learning to Predict New Views from the World's Imagery
- Generating Images Part by Part with Composite Generative Adversarial Networks
- Shape Completion Enabled Robotic Grasping