DeepStereo: Learning to Predict New Views from the World's Imagery
arXiv:1506.06825
Abstract
Deep networks have recently enjoyed enormous success when applied to recognition and classification problems in computer vision, but their use in graphics problems has been limited. In this work, we present a novel deep architecture that performs new view synthesis directly from pixels, trained from a large number of posed image sets. In contrast to traditional approaches which consist of multiple complex stages of processing, each of which require careful tuning and can fail in unexpected ways, our system is trained end-to-end. The pixels from neighboring views of a scene are presented to the network which then directly produces the pixels of the unseen view. The benefits of our approach include generality (we only require posed image sets and can easily apply our method to different domains), and high quality results on traditionally difficult scenes. We believe this is due to the end-to-end nature of our system which is able to plausibly generate pixels according to color, depth, and texture priors learnt automatically from the training data. To verify our method we show that it can convincingly reproduce known test views from nearby imagery. Additionally we show images rendered from novel viewpoints. To our knowledge, our work is the first to apply deep learning to the problem of new view synthesis from sets of real-world, natural imagery.
Video showing additional results available at http://youtu.be/cizgVZ8rjKA
References in corpus (5)
Cited by in corpus (41)
- Dynamic Filter Networks
- Deep multi-scale video prediction beyond mean square error
- Unsupervised Learning of Depth and Ego-Motion from Video
- Photographic Image Synthesis with Cascaded Refinement Networks
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Object-Centric Neural Scene Rendering
- GA-Net: Guided Aggregation Net for End-to-end Stereo Matching
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- MEMC-Net: Motion Estimation and Motion Compensation Driven Neural Network for Video Interpolation and Enhancement
- DeepMVS: Learning Multi-view Stereopsis
- Learning Dense Correspondence via 3D-guided Cycle Consistency
- AUTO3D: Novel view synthesis through unsupervisely learned variational viewpoint and global 3D representation
- View Synthesis by Appearance Flow
- DeepView: View Synthesis with Learned Gradient Descent
- 3D Photography using Context-aware Layered Depth Inpainting
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- Deep Stereo using Adaptive Thin Volume Representation with Uncertainty Awareness
- Single View Stereo Matching
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- FaceScape: a Large-scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction
- Self-Supervised Learning of Depth and Camera Motion from 360° Videos
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Dense Depth Posterior (DDP) from Single Image and Sparse Range
- LapEPI-Net: A Laplacian Pyramid EPI structure for Learning-based Dense Light Field Reconstruction
- Generative Adversarial Frontal View to Bird View Synthesis
- SegStereo: Exploiting Semantic Information for Disparity Estimation
- A Lightweight Neural Network for Monocular View Generation with Occlusion Handling
- Deep 3D-Zoom Net: Unsupervised Learning of Photo-Realistic 3D-Zoom
- Deep3D: Fully Automatic 2D-to-3D Video Conversion with Deep Convolutional Neural Networks
- LiveView: Dynamic Target-Centered MPI for View Synthesis
- Efficient Action Detection in Untrimmed Videos via Multi-Task Learning
- AIM 2019 Challenge on Video Temporal Super-Resolution: Methods and Results
- Efficient Neural Radiance Fields for Interactive Free-viewpoint Video
- Fine-scale Surface Normal Estimation using a Single NIR Image
- Incremental Scene Synthesis
- Depth Assisted Full Resolution Network for Single Image-based View Synthesis
- Multiple Kernel Learning and Automatic Subspace Relevance Determination for High-dimensional Neuroimaging Data
- Rendu basé image avec contraintes sur les gradients
- Semantic View Synthesis