activity
20232026
most citedMulti-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection

5 citations · 11 across the 26 of their papers we have counts for

collaborators
Showing cs.CVShow all

25 papers · 1 filter

cs.CV2026

Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification

Yakun Huo, Yingquan Wang, Yangyang Liu +4

RGB-Event Video Person Re-Identification (RE-VReID) aims to retrieve specific person across non-overlapping cameras with complementary RGB videos and event streams. However, existi…

cs.CV2026

Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance

Weixiang Zhou, Jiabei Zuo, Yuhao Wang +3

Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However, existing multi-modal ReID…

cs.CV2026

HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation

Pingping Zhang, Tianyu Yan, Yuhao Wang +7

Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with…

cs.CV2026

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

Hao Li, Yuhao Wang, Wenning Hao +3

RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing…

cs.CV2026

Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion

Yixin Zhu, Long Lv, Pingping Zhang +5

Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently…

cs.CV2025

X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification

Chenyang Yu, Xuehu Liu, Pingping Zhang +1

Large-scale vision-language models (e.g., CLIP) have recently achieved remarkable performance in retrieval tasks, yet their potential for Video-based Visible-Infrared Person Re-Ide…