Matterport3D: Learning from RGB-D Data in Indoor Environments
arXiv:1709.06158
Abstract
Access to large, diverse RGB-D datasets is critical for training RGB-D scene understanding algorithms. However, existing datasets still cover only a limited number of views or a restricted scale of spaces. In this paper, we introduce Matterport3D, a large-scale RGB-D dataset containing 10,800 panoramic views from 194,400 RGB-D images of 90 building-scale scenes. Annotations are provided with surface reconstructions, camera poses, and 2D and 3D semantic segmentations. The precise global alignment and comprehensive, diverse panoramic set of views over entire buildings enable a variety of supervised and self-supervised computer vision tasks, including keypoint matching, view overlap prediction, normal prediction from color, semantic segmentation, and region classification.
References in corpus (1)
Cited by in corpus (39)
- Object Goal Navigation using Goal-Oriented Semantic Exploration
- HoME: a Household Multimodal Environment
- The StreetLearn Environment and Dataset
- Benchmarking Classic and Learned Navigation in Complex 3D Environments
- VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View
- 3D-FUTURE: 3D Furniture shape with TextURE
- The Regretful Agent: Heuristic-Aided Navigation through Progress Estimation
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language Navigation
- RMM: A Recursive Mental Model for Dialog Navigation
- SceneGen: Generative Contextual Scene Augmentation using Scene Graph Priors
- Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
- Im2Pano3D: Extrapolating 360 Structure and Semantics Beyond the Field of View
- RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion
- 3D Shape Synthesis for Conceptual Design and Optimization Using Variational Autoencoders
- Simultaneous Mapping and Target Driven Navigation
- RGB-D Odometry and SLAM
- Tactical Rewind: Self-Correction via Backtracking in Vision-and-Language Navigation
- VisualEchoes: Spatial Image Representation Learning through Echolocation
- BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps
- Sim-Real Joint Reinforcement Transfer for 3D Indoor Navigation
- Efficient Semantic Scene Completion Network with Spatial Group Convolution
- Visual Question Answering on 360° Images
- Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments
- SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving
- GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes
- Deep Visual MPC-Policy Learning for Navigation
- Assisted Perception: Optimizing Observations to Communicate State
- ODE-CNN: Omnidirectional Depth Extension Networks
- Learning the Depths of Moving People by Watching Frozen People
- Restyling Data: Application to Unsupervised Domain Adaptation
- Hierarchy Denoising Recursive Autoencoders for 3D Scene Layout Prediction
- Active Imitation Learning from Multiple Non-Deterministic Teachers: Formulation, Challenges, and Algorithms
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views
- Neural Illumination: Lighting Prediction for Indoor Environments
- Associative3D: Volumetric Reconstruction from Sparse Views
- Leveraging Semantics for Incremental Learning in Multi-Relational Embeddings
- Training Deep Neural Networks to Detect Repeatable 2D Features Using Large Amounts of 3D World Capture Data