5 papers
I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs
Yu Qi, Lipeng Gu, Honghua Chen +2
Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. R…
Unified Representation Space for 3D Visual Grounding
Yinuo Zheng, Lipeng Gu, Honghua Chen +2
3D visual grounding (3DVG) is a critical task in scene understanding that aims to identify objects in 3D scenes based on text descriptions. However, existing methods rely on separa…
PointSFDA: Source-free Domain Adaptation for Point Cloud Completion
Xing He, Zhe Zhu, Liangliang Nan +3
Conventional methods for point cloud completion, typically trained on synthetic datasets, face significant challenges when applied to out-of-distribution real-world scans. In this…
CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction
Lipeng Gu, Xuefeng Yan, Weiming Wang +4
The fusion of camera- and LiDAR-based detections offers a promising solution to mitigate tracking failures in 3D multi-object tracking (MOT). However, existing methods predominantl…
PointCG: Self-supervised Point Cloud Learning via Joint Completion and Generation
Yun Liu, Peng Li, Xuefeng Yan +6
The core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects…