View Synthesis by Appearance Flow
arXiv:1605.03557
Abstract
We address the problem of novel view synthesis: given an input image, synthesizing new images of the same object or scene observed from arbitrary viewpoints. We approach this as a learning task but, critically, instead of learning to synthesize pixels from scratch, we learn to copy them from the input image. Our approach exploits the observation that the visual appearance of different views of the same instance is highly correlated, and such correlation could be explicitly learned by training a convolutional neural network (CNN) to predict appearance flows -- 2-D coordinate vectors specifying which pixels in the input view could be used to reconstruct the target view. Furthermore, the proposed framework easily generalizes to multiple input views by learning how to optimally combine single-view predictions. We show that for both objects and scenes, our approach is able to synthesize novel views of higher perceptual quality than previous CNN-based techniques.
References in corpus (7)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Deep Convolutional Inverse Graphics Network
- Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
- Discovering Hidden Factors of Variation in Deep Networks
- DeepStereo: Learning to Predict New Views from the World's Imagery
- 3D-Assisted Image Feature Synthesis for Novel Views of an Object
- Novel Views of Objects from a Single Image
Cited by in corpus (7)
- Semantic Facial Expression Editing using Autoencoded Flow
- Peeking Behind Objects: Layered Depth Prediction from a Single Image
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- 3D Shape Induction from 2D Views of Multiple Objects
- IterGANs: Iterative GANs to Learn and Control 3D Object Transformation
- 2D LiDAR Map Prediction via Estimating Motion Flow with GRU
- Predicting Ground-Level Scene Layout from Aerial Imagery