11 papers
Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification
Yakun Huo, Yingquan Wang, Yangyang Liu +4
RGB-Event Video Person Re-Identification (RE-VReID) aims to retrieve specific person across non-overlapping cameras with complementary RGB videos and event streams. However, existi…
SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
Yuhao Wang, Xiang Hu, Lixin Wang +2
Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on designing discriminative models…
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
Hao Li, Yuhao Wang, Wenning Hao +3
RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing…
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
Yixin Zhu, Long Lv, Pingping Zhang +5
Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently…
X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification
Chenyang Yu, Xuehu Liu, Pingping Zhang +1
Large-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Ide…
Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
Hui Sun, Long Lv, Pingping Zhang +4
Multi-Modal Image Fusion (MMIF) aims to integrate complementary image information from different modalities to produce informative images. Previous deep learning-based MMIF methods…