Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object Detection
arXiv:2211.16779 · doi:10.1609/aaai.v37i3.25391
Abstract
Monocular 3D object detection is a low-cost but challenging task, as it requires generating accurate 3D localization solely from a single image input. Recent developed depth-assisted methods show promising results by using explicit depth maps as intermediate features, which are either precomputed by monocular depth estimation networks or jointly evaluated with 3D object detection. However, inevitable errors from estimated depth priors may lead to misaligned semantic information and 3D localization, hence resulting in feature smearing and suboptimal predictions. To mitigate this issue, we propose ADD, an Attention-based Depth knowledge Distillation framework with 3D-aware positional encoding. Unlike previous knowledge distillation frameworks that adopt stereo- or LiDAR-based teachers, we build up our teacher with identical architecture as the student but with extra ground-truth depth as input. Credit to our teacher design, our framework is seamless, domain-gap free, easily implementable, and is compatible with object-wise ground-truth depth. Specifically, we leverage intermediate features and responses for knowledge distillation. Considering long-range 3D dependencies, we propose \emph{3D-aware self-attention} and \emph{target-aware cross-attention} modules for student adaptation. Extensive experiments are performed to verify the effectiveness of our framework on the challenging KITTI 3D object detection benchmark. We implement our framework on three representative monocular detectors, and we achieve state-of-the-art performance with no additional inference computational cost relative to baseline models. Our code is available at https://github.com/rockywind/ADD.
Accepted by AAAI2023
References in corpus (22)
- Distilling the Knowledge in a Neural Network
- FitNets: Hints for Thin Deep Nets
- MonoDistill: Learning Spatial Features for Monocular 3D Object Detection
- RTM3D: Real-time Monocular 3D Detection from Object Keypoints for Autonomous Driving
- Stereo R-CNN based 3D Object Detection for Autonomous Driving
- Learning Depth-Guided Convolutions for Monocular 3D Object Detection
- Distilling Object Detectors via Decoupled Features
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images
- Delving into Localization Errors for Monocular 3D Object Detection
- Geometry Uncertainty Projection Network for Monocular 3D Object Detection
- Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation
- Objects are Different: Flexible Monocular 3D Object Detection
- Depth-conditioned Dynamic Message Propagation for Monocular 3D Object Detection
- MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer
- MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty Propagation
- Boosting Light-Weight Depth Estimation Via Knowledge Distillation
- LIGA-Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D Detector
- Pseudo-Stereo for Monocular 3D Object Detection in Autonomous Driving
- GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection
- MonoJSG: Joint Semantic and Geometric Cost Volume for Monocular 3D Object Detection
- Monocular 3D Object Detection: An Extrinsic Parameter Free Approach
- PETR: Position Embedding Transformation for Multi-View 3D Object Detection