BundleFusion: Real-time Globally Consistent 3D Reconstruction using On-the-fly Surface Re-integration
arXiv:1604.01093
Abstract
Real-time, high-quality, 3D scanning of large-scale scenes is key to mixed reality and robotic applications. However, scalability brings challenges of drift in pose estimation, introducing significant errors in the accumulated model. Approaches often require hours of offline processing to globally correct model errors. Recent online methods demonstrate compelling results, but suffer from: (1) needing minutes to perform online correction preventing true real-time use; (2) brittle frame-to-frame (or frame-to-model) pose estimation resulting in many tracking failures; or (3) supporting only unstructured point-based representations, which limit scan quality and applicability. We systematically address these issues with a novel, real-time, end-to-end reconstruction framework. At its core is a robust pose estimation strategy, optimizing per frame for a global set of camera poses by considering the complete history of RGB-D input with an efficient hierarchical approach. We remove the heavy reliance on temporal tracking, and continually localize to the globally optimized frames instead. We contribute a parallelizable optimization framework, which employs correspondences based on sparse features and dense geometric and photometric matching. Our approach estimates globally optimized (i.e., bundle adjusted) poses in real-time, supports robust tracking with recovery from gross tracking failures (i.e., relocalization), and re-estimates the 3D model in real-time to ensure global consistency; all within a single framework. Our approach outperforms state-of-the-art online systems with quality on par to offline methods, but with unprecedented speed and scan completeness. Our framework leads to a comprehensive online scanning solution for large indoor environments, enabling ease of use and high-quality results.
References in corpus (1)
Cited by in corpus (15)
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Vox-Fusion: Dense Tracking and Mapping with Voxel-based Neural Implicit Representation
- Reference Pose Generation for Long-term Visual Localization via Learned Features and View Synthesis
- SurfelMeshing: Online Surfel-Based Mesh Reconstruction
- SLAMCast: Large-Scale, Real-Time 3D Reconstruction and Streaming for Immersive Multi-Client Live Telepresence
- Shape Completion using 3D-Encoder-Predictor CNNs and Shape Synthesis
- cilantro: A Lean, Versatile, and Efficient Library for Point Cloud Data Processing
- Opt: A Domain Specific Language for Non-linear Least Squares Optimization in Graphics and Imaging
- RGB-D Odometry and SLAM
- Plan3D: Viewpoint and Trajectory Optimization for Aerial Multi-View Stereo Reconstruction
- LiveNVS: Neural View Synthesis on Live RGB-D Streams
- Rendering the Directional TSDF for Tracking and Multi-Sensor Registration with Point-To-Plane Scale ICP
- FlyCap: Markerless Motion Capture Using Multiple Autonomous Flying Cameras
- C-blox: A Scalable and Consistent TSDF-based Dense Mapping Approach
- Improving Semantic Image Segmentation via Label Fusion in Semantically Textured Meshes