Stereo CenterNet based 3D Object Detection for Autonomous Driving
arXiv:2103.11071 · doi:10.1016/j.neucom.2021.11.048
Abstract
Recently, three-dimensional (3D) detection based on stereo images has progressed remarkably; however, most advanced methods adopt anchor-based two-dimensional (2D) detection or depth estimation to address this problem. Nevertheless, high computational cost inhibits these methods from achieving real-time performance. In this study, we propose a 3D object detection method, Stereo CenterNet (SC), using geometric information in stereo imagery. SC predicts the four semantic key points of the 3D bounding box of the object in space and utilizes 2D left and right boxes, 3D dimension, orientation, and key points to restore the bounding box of the object in the 3D space. Subsequently, we adopt an improved photometric alignment module to further optimize the position of the 3D bounding box. Experiments conducted on the KITTI dataset indicate that the proposed SC exhibits the best speed-accuracy trade-off among advanced methods without using extra data.
References in corpus (4)
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- EAO-SLAM: Monocular Semi-Dense Object SLAM Based on Ensemble Data Association
- RTM3D: Real-time Monocular 3D Detection from Object Keypoints for Autonomous Driving
- Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation