activity
20222025
most citedCLIP-Driven Fine-grained Text-Image Person Re-identification

6 citations · 6 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification

Linhan Zhou, Shuang Li, Neng Dong +3

Person re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although b…

cs.CV2025

DINOv2 Driven Gait Representation Learning for Video-Based Visible-Infrared Person Re-identification

Yujie Yang, Shuang Li, Jun Ye +3

Video-based Visible-Infrared person re-identification (VVI-ReID) aims to retrieve the same pedestrian across visible and infrared modalities from video sequences. Existing methods…

cs.CV2025

Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-Identification

Neng Dong, Shuanglin Yan, Liyan Zhang +1

Visible-Infrared Person Re-Identification (VI-ReID) is a challenging task due to the large modality discrepancy between visible and infrared images, which complicates the alignment…

cs.CV2025

ShapeSpeak: Body Shape-Aware Textual Alignment for Visible-Infrared Person Re-Identification

Shuanglin Yan, Neng Dong, Shuang Li +3

Visible-Infrared Person Re-identification (VIReID) aims to match visible and infrared pedestrian images, but the modality differences and the complexity of identity features make i…

cs.CV2024

Embedding and Enriching Explicit Semantics for Visible-Infrared Person Re-Identification

Neng Dong, Shuanglin Yan, Liyan Zhang +1

Visible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods learn visual content solely from…

cs.CV20226 cited

CLIP-Driven Fine-grained Text-Image Person Re-identification

Shuanglin Yan, Neng Dong, Liyan Zhang +1

TIReID aims to retrieve the image corresponding to the given text query from a pool of candidate images. Existing methods employ prior knowledge from single-modality pre-training t…