15 papers
Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification
Yakun Huo, Yingquan Wang, Yangyang Liu +4
RGB-Event Video Person Re-Identification (RE-VReID) aims to retrieve specific person across non-overlapping cameras with complementary RGB videos and event streams. However, existi…
Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
Weixiang Zhou, Jiabei Zuo, Yuhao Wang +3
Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However, existing multi-modal ReID…
SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
Yuhao Wang, Xiang Hu, Lixin Wang +2
Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on designing discriminative models…
HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation
Pingping Zhang, Tianyu Yan, Yuhao Wang +7
Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with…
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
Hao Li, Yuhao Wang, Wenning Hao +3
RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing…
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
Yixin Zhu, Long Lv, Pingping Zhang +5
Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently…