From the 1 of 6 linked papers with an AI index.
6 papers
What Would You Click? Personalized Video Thumbnail Generation with Preference-aware Highlight Retrieval
Zhiyu He, Zecheng Zhao, Tong Chen +3
The paper proposes a two‑stage system that first retrieves user‑specific video highlights and then generates personalized video thumbnails using a diffusion model guided by visual‑…
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
Zecheng Zhao, Zhi Chen, Zi Huang +2
Text-to-Video Retrieval (TVR) is essential in video platforms. Dense retrieval with dual-modality encoders leads in accuracy, but its computation and storage scale poorly with corp…
Distributed Zero-Shot Learning for Visual Recognition
Zhi Chen, Yadan Luo, Zi Huang +3
In this paper, we propose a Distributed Zero-Shot Learning (DistZSL) framework that can fully exploit decentralized data to learn an effective model for unseen classes. Considering…
Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
Zhi Chen, Xin Yu, Xiaohui Tao +2
Vision-language models (VLMs) such as CLIP achieve zero-shot transfer across various tasks by pre-training on numerous image-text pairs. These models often benefit from using an en…
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
Zhi Chen, Zecheng Zhao, Jingcai Guo +2
Zero-shot learning (ZSL) aims to recognize unseen classes without labeled training examples by leveraging class-level semantic descriptors such as attributes. A fundamental challen…
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
Zecheng Zhao, Zhi Chen, Zi Huang +2
Text-to-Video Retrieval (TVR) aims to retrieve relevant videos based on textual queries. However, as video content evolves continuously, adapting TVR systems to new data remains a…