1 citations · 2 across the 5 of their papers we have counts for
7 papers
ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models
Haojie Ren, Songrui Luo, Lingfeng Wang +6
LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents. However, current BEV represe…
CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models
Yifu Xiong, Wenhao Yu, Jiaxuan Lin +5
Vision-Language-Action (VLA) models have become a promising paradigm for generalist robot manipulation, where visual-language representations are used to condition continuous actio…
\(X\)-evolve: Solution space evolution powered by large language models
Yi Zhai, Zhiqiang Wei, Ruohan Li +7
While combining large language models (LLMs) with evolutionary algorithms (EAs) shows promise for solving complex optimization problems, current approaches typically evolve individ…
MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited Overlap
Yilong Wu, Yifan Duan, Yuxi Chen +5
Large-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose…
Rendering-Enhanced Automatic Image-to-Point Cloud Registration for Roadside Scenes
Yu Sheng, Lu Zhang, Xingchen Li +4
Prior point cloud provides 3D environmental context, which enhances the capabilities of monocular camera in downstream vision tasks, such as 3D object detection, via data fusion. H…
CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks
Beibei Wang, Zijian Yu, Lu Zhang +8
Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems research and development. However, it…