4 citations · 4 across the 3 of their papers we have counts for
4 papers
PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval
Xiaolun Jing, Kezhao Yin, Xinxing Yang +2
With the emergence of large-scale image-text pre-training models, e.g., CLIP, text-video retrieval has experienced substantial advances in recent years. Existing best-performing me…
Text-Video Retrieval With Global-Local Contrastive Consistency Learning
Xiaolun Jing, Xinxing Yang, Genke Yang
Text-video retrieval aims to find the most semantically similar videos with given text queries. However, since videos contain more diverse content than texts, the main semantics ex…
TC-MGC: Text-Conditioned Multi-Grained Contrastive Learning for Text-Video Retrieval
Xiaolun Jing, Genke Yang, Jian Chu
Motivated by the success of coarse-grained or fine-grained contrast in text-video retrieval, there emerge multi-grained contrastive learning methods which focus on the integration…
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
Xiaolun Jing, Genke Yang, Jian Chu
CLIP4Clip model transferred from the CLIP has been the de-factor standard to solve the video clip retrieval task from frame-level input, triggering the surge of CLIP4Clip-based mod…