1 citations · 1 across the 6 of their papers we have counts for
7 papers
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
Yudi Shi, Shangzhe Di, Qirui Chen +5
Video reasoning constitutes a comprehensive assessment of a model's capabilities, as it demands robust perceptual and interpretive skills, thereby serving as a means to explore the…
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
Jingyao Li, Jingyun Wang, Molin Tan +6
Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare informati…
USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
Jianyu Wen, Jingyun Wang, Cilin Yan +3
Recently, Large Language Models (LLMs) have been widely employed in Conversational Recommender Systems (CRSs). Unlike traditional language model approaches that focus on training,…
Object-centric Video Question Answering with Visual Grounding and Referring
Haochen Wang, Qirui Chen, Cilin Yan +5
Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level compre…
Progressive Scaling Visual Object Tracking
Jack Hong, Shilin Yan, Zehao Xiao +4
In this work, we propose a progressive scaling training strategy for visual object tracking, systematically analyzing the influence of training data volume, model size, and input r…
DynaPrompt: Dynamic Test-Time Prompt Tuning
Zehao Xiao, Shilin Yan, Jack Hong +6
Test-time prompt tuning enhances zero-shot generalization of vision-language models but tends to ignore the relatedness among test samples during inference. Online test-time prompt…