6 citations · 6 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025★ 6 cited
SCMM: Calibrating Cross-modal Representations for Text-Based Person Search
Jing Liu, Donglai Wei, Yang Liu +5
Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale gallery using natural language descriptions, posing fundamental challenges in cross-modal r…
cs.CV2024
EAVL: Explicitly Align Vision and Language for Referring Image Segmentation
Yichen Yan, Xingjian He, Wenxuan Wang +2
Referring image segmentation (RIS) aims to segment an object mentioned in natural language from an image. The main challenge is text-to-pixel fine-grained correlation. In the previ…