24 citations · 24 across the 2 of their papers we have counts for
2 papers
cs.CV2026
HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval
Zequn Xie, Xin Liu, Boyun Zhang +3
The success of CLIP has driven substantial progress in text-video retrieval. However, current methods often suffer from "blind" feature interaction, where the model struggles to di…
cs.CV2024★ 24 cited
Multi-modal Attribute Prompting for Vision-Language Models
Xin Liu, Jiamin Wu, and Wenfei Yang +2
Pre-trained Vision-Language Models (VLMs), like CLIP, exhibit strong generalization ability to downstream tasks but struggle in few-shot scenarios. Existing prompting techniques pr…