Matterport3D: Learning from RGB-D Data in Indoor Environments
arXiv:1709.06158
Abstract
Access to large, diverse RGB-D datasets is critical for training RGB-D scene understanding algorithms. However, existing datasets still cover only a limited number of views or a restricted scale of spaces. In this paper, we introduce Matterport3D, a large-scale RGB-D dataset containing 10,800 panoramic views from 194,400 RGB-D images of 90 building-scale scenes. Annotations are provided with surface reconstructions, camera poses, and 2D and 3D semantic segmentations. The precise global alignment and comprehensive, diverse panoramic set of views over entire buildings enable a variety of supervised and self-supervised computer vision tasks, including keypoint matching, view overlap prediction, normal prediction from color, semantic segmentation, and region classification.
References in corpus (1)
Cited by in corpus (71)
- Object Goal Navigation using Goal-Oriented Semantic Exploration
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- HoME: a Household Multimodal Environment
- Language-guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
- The StreetLearn Environment and Dataset
- Benchmarking Classic and Learned Navigation in Complex 3D Environments
- VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View
- 3D-FUTURE: 3D Furniture shape with TextURE
- The Regretful Agent: Heuristic-Aided Navigation through Progress Estimation
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language Navigation
- RMM: A Recursive Mental Model for Dialog Navigation
- Move to See Better: Self-Improving Embodied Object Detection
- SceneGen: Generative Contextual Scene Augmentation using Scene Graph Priors
- Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
- Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
- Panoramic Depth Estimation via Supervised and Unsupervised Learning in Indoor Scenes
- Im2Pano3D: Extrapolating 360 Structure and Semantics Beyond the Field of View
- Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
- 3D Shape Synthesis for Conceptual Design and Optimization Using Variational Autoencoders
- RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion
- Simultaneous Mapping and Target Driven Navigation
- Tactical Rewind: Self-Correction via Backtracking in Vision-and-Language Navigation
- RGB-D Odometry and SLAM
- VisualEchoes: Spatial Image Representation Learning through Echolocation
- The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark for Physically Realistic Embodied AI
- BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps
- Sim-Real Joint Reinforcement Transfer for 3D Indoor Navigation
- Efficient Semantic Scene Completion Network with Spatial Group Convolution
- Visual Question Answering on 360° Images
- An A* Curriculum Approach to Reinforcement Learning for RGBD Indoor Robot Navigation
- Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments
- SGoLAM: Simultaneous Goal Localization and Mapping for Multi-Object Goal Navigation
- SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving
- Predicting Performance of SLAM Algorithms
- NeSF: Neural Semantic Fields for Generalizable Semantic Segmentation of 3D Scenes
- SSCNav: Confidence-Aware Semantic Scene Completion for Visual Semantic Navigation
- Dynamic Plane Convolutional Occupancy Networks
- GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes
- Deep Visual MPC-Policy Learning for Navigation
- MaAST: Map Attention with Semantic Transformersfor Efficient Visual Navigation
- Restyling Data: Application to Unsupervised Domain Adaptation
- ODE-CNN: Omnidirectional Depth Extension Networks
- Assisted Perception: Optimizing Observations to Communicate State
- Learning the Depths of Moving People by Watching Frozen People
- BEyond observation: an approach for ObjectNav
- Visual Perception Generalization for Vision-and-Language Navigation via Meta-Learning
- Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation
- PanGEA: The Panoramic Graph Environment Annotation Toolkit
- Active Imitation Learning from Multiple Non-Deterministic Teachers: Formulation, Challenges, and Algorithms
- Hierarchy Denoising Recursive Autoencoders for 3D Scene Layout Prediction
- Semantic Audio-Visual Navigation
- Semantic Dense Reconstruction with Consistent Scene Segments
- The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal Navigation
- 3D Guided Weakly Supervised Semantic Segmentation
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views
- Learning Equivariant Representations
- Towards Part-Based Understanding of RGB-D Scans
- Associative3D: Volumetric Reconstruction from Sparse Views
- AccSS3D: Accelerator for Spatially Sparse 3D DNNs
- Training Deep Neural Networks to Detect Repeatable 2D Features Using Large Amounts of 3D World Capture Data
- Localising In Complex Scenes Using Balanced Adversarial Adaptation
- Learning to Explore by Reinforcement over High-Level Options
- Natural Language for Human-Robot Collaboration: Problems Beyond Language Grounding
- Leveraging Semantics for Incremental Learning in Multi-Relational Embeddings
- ADeLA: Automatic Dense Labeling with Attention for Viewpoint Adaptation in Semantic Segmentation
- Neural Illumination: Lighting Prediction for Indoor Environments
- Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
- MAOMaps: A Photo-Realistic Benchmark For vSLAM and Map Merging Quality Assessment
- VSGM -- Enhance robot task understanding ability through visual semantic graph