Monocular 3D Object Detection with Pseudo-LiDAR Point Cloud
arXiv:1903.09847
Abstract
Monocular 3D scene understanding tasks, such as object size estimation, heading angle estimation and 3D localization, is challenging. Successful modern day methods for 3D scene understanding require the use of a 3D sensor. On the other hand, single image based methods have significantly worse performance. In this work, we aim at bridging the performance gap between 3D sensing and 2D sensing for 3D object detection by enhancing LiDAR-based algorithms to work with single image input. Specifically, we perform monocular depth estimation and lift the input image to a point cloud representation, which we call pseudo-LiDAR point cloud. Then we can train a LiDAR-based 3D detection network with our pseudo-LiDAR end-to-end. Following the pipeline of two-stage 3D detection algorithms, we detect 2D object proposals in the input image and extract a point cloud frustum from the pseudo-LiDAR for each proposal. Then an oriented 3D bounding box is detected for each frustum. To handle the large amount of noise in the pseudo-LiDAR, we propose two innovations: (1) use a 2D-3D bounding box consistency constraint, adjusting the predicted 3D bounding box to have a high overlap with its corresponding 2D proposal after projecting onto the image; (2) use the instance mask instead of the bounding box as the representation of 2D proposals, in order to reduce the number of points not belonging to the object in the point cloud frustum. Through our evaluation on the KITTI benchmark, we achieve the top-ranked performance on both bird's eye view and 3D object detection among all monocular methods, effectively quadrupling the performance over previous state-of-the-art. Our code is available at https://github.com/xinshuoweng/Mono3D_PLiDAR.
Camera Ready for ICCV Workshop on "Road Scene Understanding and Autonomous Driving"
References in corpus (15)
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- HDNET: Exploiting HD Maps for 3D Object Detection
- Deep Ordinal Regression Network for Monocular Depth Estimation
- PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
- Frustum ConvNet: Sliding Frustums to Aggregate Local Point-Wise Features for Amodal 3D Object Detection
- IPOD: Intensive Point-based Object Detector for Point Cloud
- Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency
- Towards Scene Understanding with Detailed 3D Object Representations
- SPLATNet: Sparse Lattice Networks for Point Cloud Processing
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization
- Learning to Sketch with Shortcut Cycle Consistency
- PIXOR: Real-time 3D Object Detection from Point Clouds
- Multiview Supervision By Registration
- Multiview Cross-supervision for Semantic Segmentation
Cited by in corpus (7)
- Track to Reconstruct and Reconstruct to Track
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object Detection
- Single-Shot 3D Detection of Vehicles from Monocular RGB Images via Geometry Constrained Keypoints in Real-Time
- Dynamic Edge Weights in Graph Neural Networks for 3D Object Detection
- ZoomNet: Part-Aware Adaptive Zooming Neural Network for 3D Object Detection
- FusionMapping: Learning Depth Prediction with Monocular Images and 2D Laser Scans
- RoIFusion: 3D Object Detection from LiDAR and Vision