OctNetFusion: Learning Depth Fusion from Data
arXiv:1704.01047
Abstract
In this paper, we present a learning based approach to depth fusion, i.e., dense 3D reconstruction from multiple depth images. The most common approach to depth fusion is based on averaging truncated signed distance functions, which was originally proposed by Curless and Levoy in 1996. While this method is simple and provides great results, it is not able to reconstruct (partially) occluded surfaces and requires a large number frames to filter out sensor noise and outliers. Motivated by the availability of large 3D model repositories and recent advances in deep learning, we present a novel 3D CNN architecture that learns to predict an implicit surface representation from the input depth maps. Our learning based method significantly outperforms the traditional volumetric fusion approach in terms of noise reduction and outlier suppression. By learning the structure of real world 3D objects and scenes, our approach is further able to reconstruct occluded regions and to fill in gaps in the reconstruction. We demonstrate that our learning based approach outperforms both vanilla TSDF fusion as well as TV-L1 fusion on the task of volumetric fusion. Further, we demonstrate state-of-the-art 3D shape completion results.
3DV 2017, https://github.com/griegler/octnetfusion
References in corpus (4)
Cited by in corpus (16)
- Learning a Multi-View Stereo Machine
- Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction
- Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
- SPLATNet: Sparse Lattice Networks for Point Cloud Processing
- BodyNet: Volumetric Inference of 3D Human Body Shapes
- ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans
- Tangent Convolutions for Dense Prediction in 3D
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- 3D Object Reconstruction from a Single Depth View with Adversarial Learning
- Deep Generative Modeling for Scene Synthesis via Hybrid Representations
- Learning Shape Priors for Single-View 3D Completion and Reconstruction
- Local Deep Implicit Functions for 3D Shape
- 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation
- 3D Object Classification via Spherical Projections
- RayNet: Learning Volumetric 3D Reconstruction with Ray Potentials
- Scan2Mesh: From Unstructured Range Scans to 3D Meshes