Joint 2D-3D-Semantic Data for Indoor Scene Understanding
arXiv:1702.01105
Abstract
We present a dataset of large-scale indoor spaces that provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations. The dataset covers over 6,000m2 and contains over 70,000 RGB images, along with the corresponding depths, surface normals, semantic annotations, global XYZ images (all in forms of both regular and 360° equirectangular images) as well as camera information. It also includes registered raw and semantically annotated 3D meshes and point clouds. The dataset enables development of joint and cross-modal learning models and potentially unsupervised approaches utilizing the regularities present in large-scale indoor spaces. The dataset is available here: http://3Dsemantics.stanford.edu/
The dataset is available http://3Dsemantics.stanford.edu/
References in corpus (2)
Cited by in corpus (151)
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Point-Voxel CNN for Efficient 3D Deep Learning
- Learning Semantic Segmentation of Large-Scale Point Clouds with Random Sampling
- DeepGCNs: Making GCNs Go as Deep as CNNs
- Automatic reconstruction of fully volumetric 3D building models from point clouds
- Gauge Equivariant Convolutional Networks and the Icosahedral CNN
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
- TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline
- Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
- DeepGCNs: Can GCNs Go as Deep as CNNs?
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- RGB-D And Thermal Sensor Fusion: A Systematic Literature Review
- MASC: Multi-scale Affinity with Sparse Convolution for 3D Instance Segmentation
- Kaolin: A PyTorch Library for Accelerating 3D Deep Learning Research
- PVNAS: 3D Neural Architecture Search with Point-Voxel Convolution
- Pano3D: A Holistic Benchmark and a Solid Baseline for Depth Estimation
- 3D shape sensing and deep learning-based segmentation of strawberries
- A Comprehensive Review of Modern Object Segmentation Approaches
- Learning Physical Graph Representations from Visual Scenes
- Distortion-aware Monocular Depth Estimation for Omnidirectional Images
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- PCN: Point Completion Network
- H-CNN: Spatial Hashing Based CNN for 3D Shape Analysis
- 3D Point Cloud Processing and Learning for Autonomous Driving
- BiFuse++: Self-supervised and Efficient Bi-projection Fusion for 360 Depth Estimation
- Extending Maps with Semantic and Contextual Object Information for Robot Navigation: a Learning-Based Framework using Visual and Depth Cues
- 3D Scene Geometry Estimation from 360 Imagery: A Survey
- LayoutNet: Reconstructing the 3D Room Layout from a Single RGB Image
- Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-supervised Learning
- Image Amodal Completion: A Survey
- Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
- Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling
- 3D Point Cloud Descriptors in Hand-crafted and Deep Learning Age: State-of-the-Art
- HoliCity: A City-Scale Data Platform for Learning Holistic 3D Structures
- PoinTr: Diverse Point Cloud Completion with Geometry-Aware Transformers
- The AdobeIndoorNav Dataset: Towards Deep Reinforcement Learning based Real-world Indoor Robot Visual Navigation
- Semantic Segmentation for Real Point Cloud Scenes via Bilateral Augmentation and Adaptive Fusion
- Large-Field Contextual Feature Learning for Glass Detection
- Multi-source Domain Adaptation for Panoramic Semantic Segmentation
- Embodied Visual Recognition
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- A Behavioral Approach to Visual Navigation with Graph Localization Networks
- A Closer Look at Local Aggregation Operators in Point Cloud Analysis
- Depth-aware CNN for RGB-D Segmentation
- HorizonNet: Learning Room Layout with 1D Representation and Pano Stretch Data Augmentation
- Flex-Convolution (Million-Scale Point-Cloud Learning Beyond Grid-Worlds)
- GSIP: Green Semantic Segmentation of Large-Scale Indoor Point Clouds
- Surround-view Fisheye BEV-Perception for Valet Parking: Dataset, Baseline and Distortion-insensitive Multi-task Framework
- 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTics
- Neural Contourlet Network for Monocular 360 Depth Estimation
- Net: Accurate Panorama Depth Estimation on Spherical Surface
- Juggling With Representations: On the Information Transfer Between Imagery, Point Clouds, and Meshes for Multi-Modal Semantics
- TUM-FAÇADE: Reviewing and enriching point cloud benchmarks for façade segmentation
- Deep Learning for Embodied Vision Navigation: A Survey
- Panoramic Depth Estimation via Supervised and Unsupervised Learning in Indoor Scenes
- Distortion-Tolerant Monocular Depth Estimation On Omnidirectional Images Using Dual-cubemap
- PU-Transformer: Point Cloud Upsampling Transformer
- One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation
- End-to-End Wireframe Parsing
- Spin-Weighted Spherical CNNs
- SpherePHD: Applying CNNs on a Spherical PolyHeDron Representation of 360 degree Images
- Hypergraph Convolutional Network based Weakly Supervised Point Cloud Semantic Segmentation with Scene-Level Annotations
- Grid-GCN for Fast and Scalable Point Cloud Learning
- Benchmarking Deep Learning Architectures for Urban Vegetation Point Cloud Semantic Segmentation from MLS
- Efficient Urban-scale Point Clouds Segmentation with BEV Projection
- SceneCode: Monocular Dense Semantic Reconstruction using Learned Encoded Scene Representations
- Dynamics-aware Adversarial Attack of Adaptive Neural Networks
- PanoRoom: From the Sphere to the 3D Layout
- FreDSNet: Joint Monocular Depth and Semantic Segmentation with Fast Fourier Convolutions
- Discovering Closed-Loop Failures of Vision-Based Controllers via Reachability Analysis
- Spherical View Synthesis for Self-Supervised 360 Depth Estimation
- Manhattan Room Layout Reconstruction from a Single 360 image: A Comparative Study of State-of-the-art Methods
- Alias-Free Convnets: Fractional Shift Invariance via Polynomial Activations
- LayoutMP3D: Layout Annotation of Matterport3D
- Bounding Box Disparity: 3D Metrics for Object Detection With Full Degree of Freedom
- HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features
- 360SD-Net: 360° Stereo Depth Estimation with Learnable Cost Volume
- SoundSpaces: Audio-Visual Navigation in 3D Environments
- Weakly Supervised Silhouette-based Semantic Scene Change Detection
- Orientation-aware Semantic Segmentation on Icosahedron Spheres
- Point Clouds Learning with Attention-based Graph Convolution Networks
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- Multi-Modal Attention-based Fusion Model for Semantic Segmentation of RGB-Depth Images
- DRINet: A Dual-Representation Iterative Learning Network for Point Cloud Segmentation
- DuLa-Net: A Dual-Projection Network for Estimating Room Layouts from a Single RGB Panorama
- Synthetic Data Generation and Adaption for Object Detection in Smart Vending Machines
- CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth
- Atlanta Scaled layouts from non-central panoramas
- Visual Question Answering on 360° Images
- 3D-to-2D Distillation for Indoor Scene Parsing
- TextureNet: Consistent Local Parametrizations for Learning from High-Resolution Signals on Meshes
- Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation
- Learning to Recover 3D Scene Shape from a Single Image
- SILVR: A Synthetic Immersive Large-Volume Plenoptic Dataset
- Multi-Resolution Graph Neural Network for Large-Scale Pointcloud Segmentation
- LED2-Net: Monocular 360 Layout Estimation via Differentiable Depth Rendering
- Lightweight integration of 3D features to improve 2D image segmentation
- 3D Spatial Recognition without Spatially Labeled 3D
- Vision-Dialog Navigation by Exploring Cross-modal Memory
- PointAR: Efficient Lighting Estimation for Mobile Augmented Reality
- 360 Layout Estimation via Orthogonal Planes Disentanglement and Multi-view Geometric Consistency Perception
- An Evaluation of RGB and LiDAR Fusion for Semantic Segmentation
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
- Pointwise Attention-Based Atrous Convolutional Neural Networks
- Mapped Convolutions
- Bonn Activity Maps: Dataset Description
- Multi-view PointNet for 3D Scene Understanding
- Bridging Scene Understanding and Task Execution with Flexible Simulation Environments
- Landmark Policy Optimization for Object Navigation Task
- Monocular Spherical Depth Estimation with Explicitly Connected Weak Layout Cues
- Scan2Part: Fine-grained and Hierarchical Part-level Understanding of Real-World 3D Scans
- Hausdorff Point Convolution with Geometric Priors
- Layout-Guided Novel View Synthesis from a Single Indoor Panorama
- PnP-3D: A Plug-and-Play for 3D Point Clouds
- Spherical Transformer: Adapting Spherical Signal to CNNs
- ODE-CNN: Omnidirectional Depth Extension Networks
- Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
- Restyling Data: Application to Unsupervised Domain Adaptation
- Deep Visual MPC-Policy Learning for Navigation
- SALA: Soft Assignment Local Aggregation for Parameter Efficient 3D Semantic Segmentation
- End-to-end learning of keypoint detection and matching for relative pose estimation
- 3D Guided Weakly Supervised Semantic Segmentation
- Region-Transformer: Self-Attention Region Based Class-Agnostic Point Cloud Segmentation
- PointShuffleNet: Learning Non-Euclidean Features with Homotopy Equivalence and Mutual Information
- Multichannel Semantic Segmentation with Unsupervised Domain Adaptation
- Equivariant Networks for Pixelized Spheres
- Multi Voxel-Point Neurons Convolution (MVPConv) for Fast and Accurate 3D Deep Learning
- EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
- DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based Optimization
- Spatial Transformer Point Convolution
- Permutation Matters: Anisotropic Convolutional Layer for Learning on Point Clouds
- Trans4Trans: Efficient Transformer for Transparent Object Segmentation to Help Visually Impaired People Navigate in the Real World
- PanoDR: Spherical Panorama Diminished Reality for Indoor Scenes
- Learning to Stylize Novel Views
- OmniLayout: Room Layout Reconstruction from Indoor Spherical Panoramas
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views
- MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis
- Surface Regression with a Hyper-Sphere Loss
- PICCOLO: Point Cloud-Centric Omnidirectional Localization
- Background-Aware 3D Point Cloud Segmentationwith Dynamic Point Feature Aggregation
- Robust 3D Scene Segmentation through Hierarchical and Learnable Part-Fusion
- Distortion Reduction for Off-Center Perspective Projection of Panoramas
- Learning Equivariant Representations
- Multiscale Graph Construction Using Non-local Cluster Features
- Exploring Deep 3D Spatial Encodings for Large-Scale 3D Scene Understanding
- Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
- VPC-Net: Completion of 3D Vehicles from MLS Point Clouds
- SSLayout360: Semi-Supervised Indoor Layout Estimation from 360-Degree Panorama