paper

Vision Foundation Model Driven Foreground-Aware Pseudo-LiDAR Generation for Monocular 3D Object Detection

arXiv:2404.09431

Abstract

Pseudo-LiDAR has become a promising paradigm for monocular 3D object detection by transforming monocular images into point cloud representations that can be processed by LiDAR-based 3D object detectors. Recent vision foundation models provide powerful geometric and semantic priors, creating new opportunities for improving the quality of pseudo-LiDAR generation. However, effectively exploiting these priors to produce reliable pseudo-LiDAR remains challenging due to inaccurate depth estimation, insufficient foreground awareness, and redundant background points.In this paper, we propose VFMM3D, a vision foundation model-driven framework for foreground-aware pseudo-LiDAR generation. VFMM3D leverages the depth prior provided by the Depth Anything Model (DAM) and the foreground prior provided by the Segment Anything Model (SAM). A foreground-aware pseudo-LiDAR painting operation is introduced to associate projected 3D points with object-level foreground information, thereby strengthening object structures and suppressing irrelevant background regions. Furthermore, a sparsification strategy is developed to remove redundant points, reduce computational overhead, and improve the compatibility of the generated point clouds with LiDAR-based 3D object detectors. Comprehensive experiments are conducted on two challenging 3D object detection datasets, KITTI and Waymo. Our VFMM3D establishes a new state-of-the-art performance on both datasets. Additionally, experimental results demonstrate the generality of VFMM3D, showcasing its seamless integration into various LiDAR-based 3D object detectors.

12 pages, 4 figures

Vision Foundation Model Driven Foreground-Aware Pseudo-LiDAR Generation for Monocular 3D Object Detection · wovepaper