ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
arXiv:1702.04405
Abstract
A key requirement for leveraging supervised deep learning methods is the availability of large, labeled datasets. Unfortunately, in the context of RGB-D scene understanding, very little data is available -- current datasets cover a small range of scene views and have limited semantic annotations. To address this issue, we introduce ScanNet, an RGB-D video dataset containing 2.5M views in 1513 scenes annotated with 3D camera poses, surface reconstructions, and semantic segmentations. To collect this data, we designed an easy-to-use and scalable RGB-D capture system that includes automated surface reconstruction and crowdsourced semantic annotation. We show that using this data helps achieve state-of-the-art performance on several 3D scene understanding tasks, including 3D object classification, semantic voxel labeling, and CAD model retrieval. The dataset is freely available at http://www.scan-net.org.
References in corpus (4)
Cited by in corpus (211)
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Cross-view Semantic Segmentation for Sensing Surroundings
- DeepV2D: Video to Depth with Differentiable Structure from Motion
- JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds
- PointCNN: Convolution On -Transformed Points
- Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
- 3-D Scene Graph: A Sparse and Semantic Representation of Physical Environments for Intelligent Agents
- DeepGCNs: Can GCNs Go as Deep as CNNs?
- PointConv: Deep Convolutional Networks on 3D Point Clouds
- PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
- InteriorNet: Mega-scale Multi-sensor Photo-realistic Indoor Scenes Dataset
- PointRNN: Point Recurrent Neural Network for Moving Point Cloud Processing
- MASC: Multi-scale Affinity with Sparse Convolution for 3D Instance Segmentation
- SegGroup: Seg-Level Supervision for 3D Instance and Semantic Segmentation
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences
- Kaolin: A PyTorch Library for Accelerating 3D Deep Learning Research
- Peeking Behind Objects: Layered Depth Prediction from a Single Image
- 3D-BEVIS: Bird's-Eye-View Instance Segmentation
- Recurrent Slice Networks for 3D Segmentation of Point Clouds
- StructureNet: Hierarchical Graph Networks for 3D Shape Generation
- STD: Sparse-to-Dense 3D Object Detector for Point Cloud
- ShellNet: Efficient Point Cloud Convolutional Neural Networks using Concentric Shells Statistics
- Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling
- The StreetLearn Environment and Dataset
- PointASNL: Robust Point Clouds Processing using Nonlocal Neural Networks with Adaptive Sampling
- LatentGNN: Learning Efficient Non-local Relations for Visual Recognition
- Three-Filters-to-Normal: An Accurate and Ultrafast Surface Normal Estimator
- Deep Learning Based 3D Segmentation: A Survey
- 3D shape sensing and deep learning-based segmentation of strawberries
- Mining Point Cloud Local Structures by Kernel Correlation and Graph Pooling
- Addressing Overfitting on Pointcloud Classification using Atrous XCRF
- ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans
- Tangent Convolutions for Dense Prediction in 3D
- Iterative Transformer Network for 3D Point Cloud
- GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud
- Learning Gaussian Instance Segmentation in Point Clouds
- PCN: Point Completion Network
- Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation
- PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation
- Self-Supervised Learning for Domain Adaptation on Point-Clouds
- Co-Planar Parametrization for Stereo-SLAM and Visual-Inertial Odometry
- TransformerFusion: Monocular RGB Scene Reconstruction using Transformers
- MLCVNet: Multi-Level Context VoteNet for 3D Object Detection
- Deep Learning for 3D Point Cloud Understanding: A Survey
- Noise-resistant Deep Learning for Object Classification in 3D Point Clouds Using a Point Pair Descriptor
- Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
- Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling
- Deep Robust Single Image Depth Estimation Neural Network Using Scene Understanding
- OctNetFusion: Learning Depth Fusion from Data
- LSANet: Feature Learning on Point Sets by Local Spatial Aware Layer
- Shape Inpainting using 3D Generative Adversarial Network and Recurrent Convolutional Networks
- The AdobeIndoorNav Dataset: Towards Deep Reinforcement Learning based Real-world Indoor Robot Visual Navigation
- Scan2CAD: Learning CAD Model Alignment in RGB-D Scans
- Deep Depth Completion of a Single RGB-D Image
- Fully-Convolutional Point Networks for Large-Scale Point Clouds
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion
- MortonNet: Self-Supervised Learning of Local Features in 3D Point Clouds
- Shallow2Deep: Indoor Scene Modeling by Single Image Understanding
- H3DNet: 3D Object Detection Using Hybrid Geometric Primitives
- FPConv: Learning Local Flattening for Point Convolution
- Point Attention Network for Semantic Segmentation of 3D Point Clouds
- PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
- Deep attention-based classification network for robust depth prediction
- Spherical Kernel for Efficient Graph Convolution on 3D Point Clouds
- Multi-Path Region Mining For Weakly Supervised 3D Semantic Segmentation on Point Clouds
- SemanticPOSS: A Point Cloud Dataset with Large Quantity of Dynamic Instances
- PlaneNet: Piece-wise Planar Reconstruction from a Single RGB Image
- NeuralBlox: Real-Time Neural Representation Fusion for Robust Volumetric Mapping
- Automatic Generation of Constrained Furniture Layouts
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- Compositional Prototype Network with Multi-view Comparision for Few-Shot Point Cloud Semantic Segmentation
- Spatial Semantic Embedding Network: Fast 3D Instance Segmentation with Deep Metric Learning
- PU-Transformer: Point Cloud Upsampling Transformer
- Neural Sparse Voxel Fields
- Convolutional Occupancy Networks
- PU-GAN: a Point Cloud Upsampling Adversarial Network
- VoxelContext-Net: An Octree based Framework for Point Cloud Compression
- SESS: Self-Ensembling Semi-Supervised 3D Object Detection
- RealMonoDepth: Self-Supervised Monocular Depth Estimation for General Scenes
- SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans
- A Learnable Self-supervised Task for Unsupervised Domain Adaptation on Point Clouds
- ELLIPSDF: Joint Object Pose and Shape Optimization with a Bi-level Ellipsoid and Signed Distance Function Description
- Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
- To Learn or Not to Learn: Visual Localization from Essential Matrices
- Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments
- Deep Point Cloud Reconstruction
- 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection
- LiDAR-based Panoptic Segmentation via Dynamic Shifting Network
- Graph Signal Processing for Geometric Data and Beyond: Theory and Applications
- DyCo3D: Robust Instance Segmentation of 3D Point Clouds through Dynamic Convolution
- Efficient Urban-scale Point Clouds Segmentation with BEV Projection
- Real-time Progressive 3D Semantic Segmentation for Indoor Scene
- Grid-GCN for Fast and Scalable Point Cloud Learning
- DeepPerimeter: Indoor Boundary Estimation from Posed Monocular Sequences
- Adversarial Texture Optimization from RGB-D Scans
- 3D Object Detection with Pointformer
- LabelFusion: A Pipeline for Generating Ground Truth Labels for Real RGBD Data of Cluttered Scenes
- Consistent Video Depth Estimation
- Point2Node: Correlation Learning of Dynamic-Node for Point Cloud Feature Modeling
- RfD-Net: Point Scene Understanding by Semantic Instance Reconstruction
- Fast acoustic scattering using convolutional neural networks
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- Rapid Pose Label Generation through Sparse Representation of Unknown Objects
- Floor-SP: Inverse CAD for Floorplans by Sequential Room-wise Shortest Path
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual Grounding
- Real-time 3D object proposal generation and classification under limited processing resources
- RevealNet: Seeing Behind Objects in RGB-D Scans
- CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth
- 3DRM:Pair-wise relation module for 3D object detection
- Global Context Aware Convolutions for 3D Point Cloud Understanding
- OccuSeg: Occupancy-aware 3D Instance Segmentation
- 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation
- TextureNet: Consistent Local Parametrizations for Learning from High-Resolution Signals on Meshes
- Spatial-Temporal Transformer for 3D Point Cloud Sequences
- Learning Feature Descriptors using Camera Pose Supervision
- Are We Hungry for 3D LiDAR Data for Semantic Segmentation? A Survey and Experimental Study
- VIN: Voxel-based Implicit Network for Joint 3D Object Detection and Segmentation for Lidars
- RobustPointSet: A Dataset for Benchmarking Robustness of Point Cloud Classifiers
- NeuralRecon: Real-Time Coherent 3D Reconstruction from Monocular Video
- Multiview Based 3D Scene Understanding On Partial Point Sets
- A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes
- Surface Normal Estimation of Tilted Images via Spatial Rectifier
- Deep Surface Normal Estimation with Hierarchical RGB-D Fusion
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- What can I do here? Leveraging Deep 3D saliency and geometry for fast and scalable multiple affordance detection
- Dynamic Plane Convolutional Occupancy Networks
- PlaneSegNet: Fast and Robust Plane Estimation Using a Single-stage Instance Segmentation CNN
- Learning Camera Localization via Dense Scene Matching
- Investigating Attention Mechanism in 3D Point Cloud Object Detection
- Towards Panoptic 3D Parsing for Single Image in the Wild
- Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints
- Multi-view PointNet for 3D Scene Understanding
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
- Learning Transformation Synchronization
- Unsupervised Learning of Intrinsic Structural Representation Points
- DI-Fusion: Online Implicit 3D Reconstruction with Deep Priors
- A Survey on Deep Geometry Learning: From a Representation Perspective
- Normal Assisted Stereo Depth Estimation
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks
- Robust Consistent Video Depth Estimation
- Loop Closure Detection with RGB-D Feature Pyramid Siamese Networks
- Generation For Adaption: A GAN-Based Approach for 3D Domain Adaption with Point Cloud Data
- Simple multi-dataset detection
- Purely Geometric Scene Association and Retrieval - A Case for Macro Scale 3D Geometry
- 3D Brain Reconstruction by Hierarchical Shape-Perception Network from a Single Incomplete Image
- Refer-it-in-RGBD: A Bottom-up Approach for 3D Visual Grounding in RGBD Images
- Improving Monocular Depth Estimation by Leveraging Structural Awareness and Complementary Datasets
- Domain Adaptation of Networks for Camera Pose Estimation: Learning Camera Pose Estimation Without Pose Labels
- ODE-CNN: Omnidirectional Depth Extension Networks
- Learning the Depths of Moving People by Watching Frozen People
- Scan2Mesh: From Unstructured Range Scans to 3D Meshes
- Analysis and Modeling of 3D Indoor Scenes
- Improving the generalization of network based relative pose regression: dimension reduction as a regularizer
- I3DOL: Incremental 3D Object Learning without Catastrophic Forgetting
- Learning to Optimize Non-Rigid Tracking
- Multimodal Semantic Scene Graphs for Holistic Modeling of Surgical Procedures
- PointMixer: MLP-Mixer for Point Cloud Understanding
- Hierarchy Denoising Recursive Autoencoders for 3D Scene Layout Prediction
- Single-Image Piece-wise Planar 3D Reconstruction via Associative Embedding
- Multichannel Semantic Segmentation with Unsupervised Domain Adaptation
- SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
- Improving Semantic Segmentation through Spatio-Temporal Consistency Learned from Videos
- ParaNet: Deep Regular Representation for 3D Point Clouds
- SALA: Soft Assignment Local Aggregation for Parameter Efficient 3D Semantic Segmentation
- MOLTR: Multiple Object Localisation, Tracking, and Reconstruction from Monocular RGB Videos
- 3D Guided Weakly Supervised Semantic Segmentation
- SLAM in the Field: An Evaluation of Monocular Mapping and Localization on Challenging Dynamic Agricultural Environment
- Deep Multi-view Depth Estimation with Predicted Uncertainty
- Semantic Dense Reconstruction with Consistent Scene Segments
- CodeMapping: Real-Time Dense Mapping for Sparse SLAM using Compact Scene Representations
- Point Cloud Instance Segmentation with Semi-supervised Bounding-Box Mining
- Shape from Shading through Shape Evolution
- Deeper or Wider Networks of Point Clouds with Self-attention?
- Few-shot 3D Point Cloud Semantic Segmentation
- SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings
- Going Deeper with Lean Point Networks
- Learning Single-Image Depth from Videos using Quality Assessment Networks
- Complete 3D Scene Parsing from an RGBD Image
- Sketch-and-test: picture-centered research with p5.js assisted crowdsourcing
- Torch-Points3D: A Modular Multi-Task Frameworkfor Reproducible Deep Learning on 3D Point Clouds
- Geometry-Aware Self-Training for Unsupervised Domain Adaptationon Object Point Clouds
- 3D Annotation Of Arbitrary Objects In The Wild
- Robust 3D Scene Segmentation through Hierarchical and Learnable Part-Fusion
- A Data-driven Prior on Facet Orientation for Semantic Mesh Labeling
- Non-local RoIs for Instance Segmentation
- Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
- BIDCD -- Bosch Industrial Depth Completion Dataset
- Indoor Semantic Scene Understanding using Multi-modality Fusion
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views
- ICM-3D: Instantiated Category Modeling for 3D Instance Segmentation
- Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic Segmentation
- AccSS3D: Accelerator for Spatially Sparse 3D DNNs
- Towards Part-Based Understanding of RGB-D Scans
- Learning Equivariant Representations
- Fine-Grained Vehicle Perception via 3D Part-Guided Visual Data Augmentation
- Boundary-induced and scene-aggregated network for monocular depth prediction
- Robust 2D/3D Vehicle Parsing in CVIS
- Picasso: A CUDA-based Library for Deep Learning over 3D Meshes
- Contextual Scene Augmentation and Synthesis via GSACNet
- A Front-End for Dense Monocular SLAM using a Learned Outlier Mask Prior
- MNEW: Multi-domain Neighborhood Embedding and Weighting for Sparse Point Clouds Segmentation
- Associative3D: Volumetric Reconstruction from Sparse Views
- Video Depth Estimation by Fusing Flow-to-Depth Proposals
- 3D Objectness Estimation via Bottom-up Regret Grouping
- Training Deep Neural Networks to Detect Repeatable 2D Features Using Large Amounts of 3D World Capture Data
- Path-Invariant Map Networks
- SK-Net: Deep Learning on Point Cloud via End-to-end Discovery of Spatial Keypoints