2 citations · 6 across the 5 of their papers we have counts for
5 papers
Disentangle and denoise: Tackling context misalignment for video moment retrieval
Kaijing Ma, Han Fang, Xianghao Zang +7
Video Moment Retrieval, which aims to locate in-context video moments according to a natural language query, is an essential task for cross-modal grounding. Existing methods focus…
ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
Han Fang, Xianghao Zang, Chao Ban +5
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since vid…
Integrating Listwise Ranking into Pairwise-based Image-Text Retrieval
Zheng Li, Caili Guo, Xin Wang +2
Image-Text Retrieval (ITR) is essentially a ranking problem. Given a query caption, the goal is to rank candidate images by relevance, from large to small. The current ITR datasets…
Selectively Hard Negative Mining for Alleviating Gradient Vanishing in Image-Text Matching
Zheng Li, Caili Guo, Xin Wang +2
Recently, a series of Image-Text Matching (ITM) methods achieve impressive performance. However, we observe that most existing ITM models suffer from gradients vanishing at the beg…
Unified Loss of Pair Similarity Optimization for Vision-Language Retrieval
Zheng Li, Caili Guo, Xin Wang +3
There are two popular loss functions used for vision-language retrieval, i.e., triplet loss and contrastive learning loss, both of them essentially minimize the difference between…