From the 1 of 5 linked papers with an AI index.
5 papers
What Would You Click? Personalized Video Thumbnail Generation with Preference-aware Highlight Retrieval
Zhiyu He, Zecheng Zhao, Tong Chen +3
The paper proposes a two‑stage system that first retrieves user‑specific video highlights and then generates personalized video thumbnails using a diffusion model guided by visual‑…
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
Zecheng Zhao, Zhi Chen, Zi Huang +2
Text-to-Video Retrieval (TVR) is essential in video platforms. Dense retrieval with dual-modality encoders leads in accuracy, but its computation and storage scale poorly with corp…
Learning A Universal Crime Predictor with Knowledge-guided Hypernetworks
Fidan Karimova, Tong Chen, Yu Yang +1
Predicting crimes in urban environments is crucial for public safety, yet existing prediction methods often struggle to align the knowledge across diverse cities that vary dramatic…
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
Zecheng Zhao, Selena Song, Tong Chen +3
Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synt…
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
Zecheng Zhao, Zhi Chen, Zi Huang +2
Text-to-Video Retrieval (TVR) aims to retrieve relevant videos based on textual queries. However, as video content evolves continuously, adapting TVR systems to new data remains a…