RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search
arXiv:2305.13653 · doi:10.24963/ijcai.2023/62
Abstract
Text-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including two novel tasks: Relation-Aware learning (RA) and Sensitivity-Aware learning (SA). For one thing, existing methods cluster representations of all positive pairs without distinction and overlook the noise problem caused by the weak positive pairs where the text and the paired image have noise correspondences, thus leading to overfitting learning. RA offsets the overfitting risk by introducing a novel positive relation detection task (i.e., learning to distinguish strong and weak positive pairs). For another thing, learning invariant representation under data augmentation (i.e., being insensitive to some transformations) is a general practice for improving representation's robustness in existing methods. Beyond that, we encourage the representation to perceive the sensitive transformation by SA (i.e., learning to detect the replaced words), thus promoting the representation's robustness. Experiments demonstrate that RaSa outperforms existing state-of-the-art methods by 6.94%, 4.45% and 15.35% in terms of Rank@1 on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets, respectively. Code is available at: https://github.com/Flame-Chasers/RaSa.
Accepted by IJCAI 2023. Code is available at https://github.com/Flame-Chasers/RaSa
References in corpus (13)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Bootstrap your own latent: A new approach to self-supervised Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- CAIBC: Capturing All-round Information Beyond Color for Text-based Person Retrieval
- Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification
- Person Search with Natural Language Description
- Contextual Non-Local Alignment over Full-Scale Representation for Text-Based Person Search
- Text-Based Person Search with Limited Data
- Image-text Retrieval: A Survey on Recent Research and Development
- TIPCB: A Simple but Effective Part-based Convolutional Baseline for Text-based Person Search
Cited by in corpus (4)
- MARS: Paying more attention to visual attributes for text-based person search
- From Data Deluge to Data Curation: A Filtering-WoRA Paradigm for Efficient Text-based Person Search
- Robust Duality Learning for Unsupervised Visible-Infrared Person Re-Identification
- DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search