AUTO3D: Novel view synthesis through unsupervisely learned variational viewpoint and global 3D representation
arXiv:2007.06620 · doi:10.1007/978-3-030-58545-7_4
Abstract
This paper targets on learning-based novel view synthesis from a single or limited 2D images without the pose supervision. In the viewer-centered coordinates, we construct an end-to-end trainable conditional variational framework to disentangle the unsupervisely learned relative-pose/rotation and implicit global 3D representation (shape, texture and the origin of viewer-centered coordinates, etc.). The global appearance of the 3D object is given by several appearance-describing images taken from any number of viewpoints. Our spatial correlation module extracts a global 3D representation from the appearance-describing images in a permutation invariant manner. Our system can achieve implicitly 3D understanding without explicitly 3D reconstruction. With an unsupervisely learned viewer-centered relative-pose/rotation code, the decoder can hallucinate the novel view continuously by sampling the relative-pose in a prior distribution. In various applications, we demonstrate that our model can achieve comparable or even better results than pose/3D model-supervised learning-based novel view synthesis (NVS) methods with any number of input views.
ECCV 2020
References in corpus (7)
- Disentangling factors of variation in deep representations using adversarial training
- Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction
- DeepStereo: Learning to Predict New Views from the World's Imagery
- Learning to Reconstruct Shapes from Unseen Classes
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Conservative Wasserstein Training for Pose Estimation
- Towards Disentangled Representations for Human Retargeting by Multi-view Learning
Cited by in corpus (6)
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Dual-cycle Constrained Bijective VAE-GAN For Tagged-to-Cine Magnetic Resonance Image Synthesis
- Reinforced Wasserstein Training for Severity-Aware Semantic Segmentation in Autonomous Driving
- Energy-constrained Self-training for Unsupervised Domain Adaptation
- Identity-aware Facial Expression Recognition in Compressed Video
- Novel View Synthesis from a Single Image via Unsupervised learning