pixelNeRF: Neural Radiance Fields from One or Few Images
arXiv:2012.02190
Abstract
We propose pixelNeRF, a learning framework that predicts a continuous neural scene representation conditioned on one or few input images. The existing approach for constructing neural radiance fields involves optimizing the representation to every scene independently, requiring many calibrated views and significant compute time. We take a step towards resolving these shortcomings by introducing an architecture that conditions a NeRF on image inputs in a fully convolutional manner. This allows the network to be trained across multiple scenes to learn a scene prior, enabling it to perform novel view synthesis in a feed-forward manner from a sparse set of views (as few as one). Leveraging the volume rendering approach of NeRF, our model can be trained directly from images with no explicit 3D supervision. We conduct extensive experiments on ShapeNet benchmarks for single image novel view synthesis tasks with held-out objects as well as entire unseen categories. We further demonstrate the flexibility of pixelNeRF by demonstrating it on multi-object ShapeNet scenes and real scenes from the DTU dataset. In all cases, pixelNeRF outperforms current state-of-the-art baselines for novel view synthesis and single image 3D reconstruction. For the video and code, please visit the project website: https://alexyu.net/pixelnerf
CVPR 2021
References in corpus (11)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision
- Semantic Image Synthesis with Spatially-Adaptive Normalization
- Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images
- Soft Rasterizer: A Differentiable Renderer for Image-based 3D Reasoning
- Learning to Infer Implicit Surfaces without 3D Supervision
- Learning to Reconstruct Shapes from Unseen Classes
- Equivariant Neural Rendering
- Pixels, voxels, and views: A study of shape representations for single view 3D object shape prediction
- Neural Rerendering in the Wild
Cited by in corpus (11)
- Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance Rendering
- Neural Volume Rendering: NeRF And Beyond
- NeRF-VAE: A Geometry Aware 3D Scene Generative Model
- Representing Long Volumetric Video with Temporal Gaussian Hierarchy
- GANeRF: Leveraging Discriminators to Optimize Neural Radiance Fields
- Generative Models as Distributions of Functions
- Lite2Relight: 3D-aware Single Image Portrait Relighting
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
- ReShader: View-Dependent Highlights for Single Image View-Synthesis
- Shape from Blur: Recovering Textured 3D Shape and Motion of Fast Moving Objects
- iButter: Neural Interactive Bullet Time Generator for Human Free-viewpoint Rendering