Multi-Sensor 3D Object Box Refinement for Autonomous Driving
arXiv:1909.04942
Abstract
We propose a 3D object detection system with multi-sensor refinement in the context of autonomous driving. In our framework, the monocular camera serves as the fundamental sensor for 2D object proposal and initial 3D bounding box prediction. While the stereo cameras and LiDAR are treated as adaptive plug-in sensors to refine the 3D box localization performance. For each observed element in the raw measurement domain (e.g., pixels for stereo, 3D points for LiDAR), we model the local geometry as an instance vector representation, which indicates the 3D coordinate of each element respecting to the object frame. Using this unified geometric representation, the 3D object location can be unified refined by the stereo photometric alignment or point cloud alignment. We demonstrate superior 3D detection and localization performance compared to state-of-the-art monocular, stereo methods and competitive performance compared with the baseline LiDAR method on the KITTI object benchmark.
References in corpus (9)
- Deep Continuous Fusion for Multi-Sensor 3D Object Detection
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
- Frustum ConvNet: Sliding Frustums to Aggregate Local Point-Wise Features for Amodal 3D Object Detection
- Joint 3D Proposal Generation and Object Detection from View Aggregation
- Pseudo-LiDAR from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving
- Stereo R-CNN based 3D Object Detection for Autonomous Driving
- MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization