3 papers
cs.CV2025
Adaptive Agent Selection and Interaction Network for Image-to-point cloud Registration
Zhixin Cheng, Xiaotian Yin, Jiacheng Deng +5
Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. Howev…
cs.CV2025
UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
Yinchao Ma, Yuyang Tang, Wenfei Yang +3
Single object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, R…
cs.CV2025
GRAM-MAMBA: Holistic Feature Alignment for Wireless Perception with Adaptive Low-Rank Compensation
Weiqi Yang, Xu Zhou, Jingfu Guan +2
Multi-modal fusion is crucial for Internet of Things (IoT) perception, widely deployed in smart homes, intelligent transport, industrial automation, and healthcare. However, existi…