computer vision

Structural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-Identification

arXiv:2607.15220

summary

The paper proposes a framework called Structural‑Semantic Reciprocal Learning (SSRL) that uses fine‑grained body‑part features and a closed‑loop semantic calibration to improve unsupervised visible‑infrared person re‑identification by reducing modality gap and filtering noisy pseudo‑labels.

Abstract

Unsupervised visible-infrared person re-identification (USVI-ReID) is challenging due to the large modality gap and the lack of cross-modal identity annotations. Progressive association paradigms have been proposed to gradually bridge the gap, but they suffer from two critical bottlenecks: reliance on ambiguous global representations and unchecked propagation of pseudo-label noise in an open-loop manner. To address these issues, we propose Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system. Structurally, we introduce Fine-grained Structural Decoupling (FSD) to extract discriminative body-part primitives as reliable spatial anchors, complementing ambiguous holistic silhouettes with spatially consistent structural details. Semantically, we design a Closed-loop Semantic Calibration (CSC) mechanism that reconstructs shared semantic prototypes at each epoch and feeds them back into the training loop, effectively filtering pseudo-label noise before the next clustering cycle. Through the reciprocal interaction between structural and semantic learning, SSRL achieves robust cross-modal representation. Extensive experiments demonstrate the competitive performance of SSRL against state-of-the-art USVI-ReID methods on both SYSU-MM01 and RegDB, notably surpassing several supervised counterparts on RegDB.

Accepted by PRCV 2026

Topics & keywords

#unsupervised person re-identification#visible-infrared cross‑modal matching#structural decoupling#semantic calibration#pseudo‑label noise reductionfine-grained structural decouplingclosed-loop semantic calibrationcross-modal representationclusteringpseudo-label filtering
Structural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-Identification · wovepaper