3 citations · 4 across the 4 of their papers we have counts for
4 papers
Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection
Xingyu Peng, Yan Bai, Chen Gao +5
Open-Vocabulary Detection (OVD) is the task of detecting all interesting objects in a given scene without predefined object classes. Extensive work has been done to deal with the O…
Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection
Jiahui Fu, Chen Gao, Zitian Wang +4
Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird's-eye view (BEV) representation space. However, our empirical findi…
Two-Stream Networks for Object Segmentation in Videos
Hannan Lu, Zhi Tian, Lirong Yang +2
Existing matching-based approaches perform video object segmentation (VOS) via retrieving support features from a pixel-level memory, while some pixels may suffer from lack of corr…
Target-Driven Structured Transformer Planner for Vision-Language Navigation
Yusheng Zhao, Jinyu Chen, Chen Gao +5
Vision-language navigation is the task of directing an embodied agent to navigate in 3D scenes with natural language instructions. For the agent, inferring the long-term navigation…