4 papers
AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
Kang Zeng, Guojin Zhong, Jintao Cheng +2
The advancement of Multimodal Large Language Models (MLLMs) has driven significant progress in Visual Question Answering (VQA), evolving from Single to Multi Image VQA (MVQA). Howe…
L2G-Map: Local-to-Global Mapping via Hierarchical Diffusion Refinement and Elliptical Bayesian Fusion
Siyu Li, Xinying Hong, Fei Teng +5
Offline high-definition maps provide essential geometric and topological priors for autonomous driving systems. Pure-vision solutions have become the predominant paradigm for offli…
MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model
Kang Zeng, Hao Shi, Jiacheng Lin +5
LiDAR-based Moving Object Segmentation (MOS) aims to locate and segment moving objects in point clouds of the current scan using motion information from previous scans. Despite the…
MF-MOS: A Motion-Focused Model for Moving Object Segmentation
Jintao Cheng, Kang Zeng, Zhuoxu Huang +5
Moving object segmentation (MOS) provides a reliable solution for detecting traffic participants and thus is of great interest in the autonomous driving field. Dynamic capture is a…