3D Object Proposals using Stereo Imagery for Accurate Object Class Detection
arXiv:1608.07711
Abstract
The goal of this paper is to perform 3D object detection in the context of autonomous driving. Our method first aims at generating a set of high-quality 3D object proposals by exploiting stereo imagery. We formulate the problem as minimizing an energy function that encodes object size priors, placement of objects on the ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. We then exploit a CNN on top of these proposals to perform object detection. In particular, we employ a convolutional neural net (CNN) that exploits context and depth information to jointly regress to 3D bounding box coordinates and object pose. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. When combined with the CNN, our approach outperforms all existing results in object detection and orientation estimation tasks for all three KITTI object classes. Furthermore, we experiment also with the setting where LIDAR information is available, and show that using both LIDAR and stereo leads to the best result.
14 pages, 12 figures
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- What makes for effective detection proposals?
- Towards Scene Understanding with Detailed 3D Object Representations
- segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection
- DeepProposal: Hunting Objects by Cascading Deep Convolutional Layers
- Hierarchical Adaptive Structural SVM for Domain Adaptation
Cited by in corpus (12)
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- Multi-View 3D Object Detection Network for Autonomous Driving
- Stereo R-CNN based 3D Object Detection for Autonomous Driving
- BEV-Seg: Bird's Eye View Semantic Segmentation Using Geometry and Semantic Point Cloud
- Locating 3D Object Proposals: A Depth-Based Online Approach
- DSGN: Deep Stereo Geometry Network for 3D Object Detection
- Driving Datasets Literature Review
- What You See is What You Get: Exploiting Visibility for 3D Object Detection
- Multi-Sensor 3D Object Box Refinement for Autonomous Driving
- Scene Understanding Networks for Autonomous Driving based on Around View Monitoring System
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Exploring intermediate representation for monocular vehicle pose estimation