most citedFine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

29 citations · 29 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2025

HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval

Zhiwei Chen, Yupeng Hu, Zixu Li +3

Composed Video Retrieval (CVR) is a challenging video retrieval task that utilizes multi-modal queries, consisting of a reference video and modification text, to retrieve the desir…

cs.CL2025

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen +3

Textual response generation is pivotal for multimodal \mbox{task-oriented} dialog systems, which aims to generate proper textual responses based on the multimodal context. While ex…

cs.CV2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

Haoyu Zhang, Meng Liu, Zaijing Li +4

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied in…

cs.CV2025

FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval

Zixu Li, Zhiheng Fu, Yupeng Hu +3

Composed Image Retrieval (CIR) facilitates image retrieval through a multimodal query consisting of a reference image and modification text. The reference image defines the retriev…

cs.CV202529 cited

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Haoqiang Lin, Haokun Wen, Xuemeng Song +3

Composed Image Retrieval (CIR) allows users to search target images with a multimodal query, comprising a reference image and a modification text that describes the user's modifica…

cs.MM2025

A Comprehensive Survey on Composed Image Retrieval

Xuemeng Song, Haoqiang Lin, Haokun Wen +3

Composed Image Retrieval (CIR) is an emerging yet challenging task that allows users to search for target images using a multimodal query, comprising a reference image and a modifi…