11 citations · 16 across the 3 of their papers we have counts for
4 papers
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
Xiangchen Liu, Hanghan Zheng, Jeil Jeong +5
Vision-language Navigation (VLN) requires an agent to understand visual observations and language instructions to navigate in unseen environments. Most existing approaches rely on…
FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection
Chunyong Hu, Hang Zheng, Kun Li +11
Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features in…
FusionAD: Multi-modality Fusion for Prediction and Planning Tasks of Autonomous Driving
Tengju Ye, Wei Jing, Chunyong Hu +11
Building a multi-modality multi-task neural network toward accurate and robust performance is a de-facto standard in perception task of autonomous driving. However, leveraging such…
CMDFusion: Bidirectional Fusion Network with Cross-modality Knowledge Distillation for LIDAR Semantic Segmentation
Jun Cen, Shiwei Zhang, Yixuan Pei +5
2D RGB images and 3D LIDAR point clouds provide complementary knowledge for the perception system of autonomous vehicles. Several 2D and 3D fusion methods have been explored for th…