Monocular 3D Object Detection: An Extrinsic Parameter Free Approach
arXiv:2106.15796
Abstract
Monocular 3D object detection is an important task in autonomous driving. It can be easily intractable where there exists ego-car pose change w.r.t. ground plane. This is common due to the slight fluctuation of road smoothness and slope. Due to the lack of insight in industrial application, existing methods on open datasets neglect the camera pose information, which inevitably results in the detector being susceptible to camera extrinsic parameters. The perturbation of objects is very popular in most autonomous driving cases for industrial products. To this end, we propose a novel method to capture camera pose to formulate the detector free from extrinsic perturbation. Specifically, the proposed framework predicts camera extrinsic parameters by detecting vanishing point and horizon change. A converter is designed to rectify perturbative features in the latent space. By doing so, our 3D detector works independent of the extrinsic parameter variations and produces accurate results in realistic cases, e.g., potholed and uneven roads, where almost all existing monocular detectors fail to handle. Experiments demonstrate our method yields the best performance compared with the other state-of-the-arts by a large margin on both KITTI 3D and nuScenes datasets.
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- DeepVO: Towards End-to-End Visual Odometry with Deep Recurrent Convolutional Neural Networks
- Deep Continuous Fusion for Multi-Sensor 3D Object Detection
- RefinedMPL: Refined Monocular PseudoLiDAR for 3D Object Detection in Autonomous Driving
- Recurrent Neural Network for (Un-)supervised Learning of Monocular VideoVisual Odometry and Depth
- Beyond Tracking: Selecting Memory and Refining Poses for Deep Visual Odometry
- Center3D: Center-based Monocular 3D Object Detection with Joint Depth Understanding
- Reconstructing Vechicles from a Single Image: Shape Priors for Road Scene Understanding
- RadarNet: Exploiting Radar for Robust Perception of Dynamic Objects
- Rethinking Pseudo-LiDAR Representation
- Deep Learning on Radar Centric 3D Object Detection
- Learning Monocular Visual Odometry via Self-Supervised Long-Term Modeling