5 papers
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
Jingyao Li, Jingyun Wang, Molin Tan +6
Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare informati…
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
Cilin Yan, Jingyun Wang, Guoliang Kang
Referring Video Segmentation (RVOS) aims to segment objects in videos given linguistic expressions. The key to solving RVOS is to extract long-range temporal context information fr…
USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
Jianyu Wen, Jingyun Wang, Cilin Yan +3
Recently, Large Language Models (LLMs) have been widely employed in Conversational Recommender Systems (CRSs). Unlike traditional language model approaches that focus on training,…
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
Huijie Liu, Jingyun Wang, Shuai Ma +3
Motion customization aims to adapt the diffusion model (DM) to generate videos with the motion specified by a set of video clips with the same motion concept. To realize this goal,…
Efficient and Accurate Prompt Optimization: the Benefit of Memory in Exemplar-Guided Reflection
Cilin Yan, Jingyun Wang, Lin Zhang +6
Automatic prompt engineering aims to enhance the generation quality of large language models (LLMs). Recent works utilize feedbacks generated from erroneous cases to guide the prom…