Showing cs.IRShow all
3 papers · 1 filter
cs.IR2026
PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval
Xiaolun Jing, Kezhao Yin, Xinxing Yang +2
With the emergence of large-scale image-text pre-training models, e.g., CLIP, text-video retrieval has experienced substantial advances in recent years. Existing best-performing me…
cs.IR2025
TC-MGC: Text-Conditioned Multi-Grained Contrastive Learning for Text-Video Retrieval
Xiaolun Jing, Genke Yang, Jian Chu
Motivated by the success of coarse-grained or fine-grained contrast in text-video retrieval, there emerge multi-grained contrastive learning methods which focus on the integration…
cs.IR2024
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
Xiaolun Jing, Genke Yang, Jian Chu
CLIP4Clip model transferred from the CLIP has been the de-factor standard to solve the video clip retrieval task from frame-level input, triggering the surge of CLIP4Clip-based mod…