Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
arXiv:1803.03352 · doi:10.1109/ACCESS.2018.2886133
Abstract
With the availability of low-cost and compact 2.5/3D visual sensing devices, computer vision community is experiencing a growing interest in visual scene understanding of indoor environments. This survey paper provides a comprehensive background to this research topic. We begin with a historical perspective, followed by popular 3D data representations and a comparative analysis of available datasets. Before delving into the application specific details, this survey provides a succinct introduction to the core technologies that are the underlying methods extensively used in the literature. Afterwards, we review the developed techniques according to a taxonomy based on the scene understanding tasks. This covers holistic indoor scene understanding as well as subtasks such as scene classification, object detection, pose estimation, semantic segmentation, 3D reconstruction, saliency detection, physics-based reasoning and affordance prediction. Later on, we summarize the performance metrics used for evaluation in different tasks and a quantitative comparison among the recent state-of-the-art techniques. We conclude this review with the current challenges and an outlook towards the open research problems requiring further investigation.
IEEE Access
References in corpus (19)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- Joint 2D-3D-Semantic Data for Indoor Scene Understanding
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Generative and Discriminative Voxel Modeling with Convolutional Neural Networks
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Distilling a Neural Network Into a Soft Decision Tree
- Adversarial Examples: Attacks and Defenses for Deep Learning
- FusionNet: 3D Object Classification Using Multiple Data Representations
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- Towards Scene Understanding with Detailed 3D Object Representations
- Shape Completion using 3D-Encoder-Predictor CNNs and Shape Synthesis
- Generic 3D Representation via Pose Estimation and Matching
- Siamese Regression Networks with Efficient mid-level Feature Extraction for 3D Object Pose Estimation
- OctNetFusion: Learning Depth Fusion from Data
- PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
Cited by in corpus (12)
- VddNet: Vine Disease Detection Network Based on Multispectral Images and Depth Map
- Deep Learning Based 3D Segmentation: A Survey
- Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
- Virtual replicas of real places: Experimental investigations
- Lightweight Residual Densely Connected Convolutional Neural Network
- FreDSNet: Joint Monocular Depth and Semantic Segmentation with Fast Fourier Convolutions
- Atlanta Scaled layouts from non-central panoramas
- Learning to Assess Danger from Movies for Cooperative Escape Planning in Hazardous Environments
- Active Scene Understanding via Online Semantic Reconstruction
- Review on 6D Object Pose Estimation with the focus on Indoor Scene Understanding
- Rethinking Data Input for Point Cloud Upsampling
- Floorplan-Jigsaw: Jointly Estimating Scene Layout and Aligning Partial Scans