Deep Sliding Shapes for Amodal 3D Object Detection in RGB-D Images
arXiv:1511.02300
Abstract
We focus on the task of amodal 3D object detection in RGB-D images, which aims to produce a 3D bounding box of an object in metric form at its full extent. We introduce Deep Sliding Shapes, a 3D ConvNet formulation that takes a 3D volumetric scene from a RGB-D image as input and outputs 3D object bounding boxes. In our approach, we propose the first 3D Region Proposal Network (RPN) to learn objectness from geometric shapes and the first joint Object Recognition Network (ORN) to extract geometric features in 3D and color features in 2D. In particular, we handle objects of various sizes by training an amodal RPN at two different scales and an ORN to regress 3D bounding boxes. Experiments show that our algorithm outperforms the state-of-the-art by 13.8 in mAP and is 200x faster than the original Sliding Shapes. All source code and pre-trained models will be available at GitHub.
Cited by in corpus (22)
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- Cross-view Semantic Segmentation for Sensing Surroundings
- From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network
- Shape Completion using 3D-Encoder-Predictor CNNs and Shape Synthesis
- Multi-View 3D Object Detection Network for Autonomous Driving
- PPFNet: Global Context Aware Local Features for Robust 3D Point Matching
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- Learning Dense Correspondence via 3D-guided Cycle Consistency
- Locating 3D Object Proposals: A Depth-Based Online Approach
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- A survey of Object Classification and Detection based on 2D/3D data
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- DeepContext: Context-Encoding Neural Pathways for 3D Holistic Scene Understanding
- Points2Pix: 3D Point-Cloud to Image Translation using conditional Generative Adversarial Networks
- RevealNet: Seeing Behind Objects in RGB-D Scans
- Monte Carlo Scene Search for 3D Scene Understanding
- Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Scan2Mesh: From Unstructured Range Scans to 3D Meshes
- LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling