5 papers
RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval
Jiale Huang, Zixu Li, Zhiheng Fu +3
Composed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. However, the prevalence of Noi…
IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval
Jiale Huang, Zixu Li, Zhiwei Chen +3
Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal cor…
INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval
Zhiwei Chen, Yupeng Hu, Zhiheng Fu +4
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modif…
MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
Zechao Zhan, Dehong Gao, Jinxia Zhang +3
Text-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate…
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
Jiale Huang, Dehong Gao, Jinxia Zhang +3
Large-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained at…