3 papers
cs.IR2026
PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval
Xiaolun Jing, Kezhao Yin, Xinxing Yang +2
With the emergence of large-scale image-text pre-training models, e.g., CLIP, text-video retrieval has experienced substantial advances in recent years. Existing best-performing me…
cs.IR2026
Text-Video Retrieval With Global-Local Contrastive Consistency Learning
Xiaolun Jing, Xinxing Yang, Genke Yang
Text-video retrieval aims to find the most semantically similar videos with given text queries. However, since videos contain more diverse content than texts, the main semantics ex…
cs.IR2025
TC-MGC: Text-Conditioned Multi-Grained Contrastive Learning for Text-Video Retrieval
Xiaolun Jing, Genke Yang, Jian Chu
Motivated by the success of coarse-grained or fine-grained contrast in text-video retrieval, there emerge multi-grained contrastive learning methods which focus on the integration…