7 papers
Order-Aware Test-Time Adaptation: Leveraging Temporal Dynamics for Robust Streaming Inference
Young Kyung Kim, Oded Schlesinger, Qiangqiang Wu +2
Test-Time Adaptation (TTA) enables pre-trained models to adjust to distribution shift by learning from unlabeled test-time streams. However, existing methods typically treat these…
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
Bo Fang, Yuxin Song, Qiangqiang Wu +3
Complex video reasoning remains a significant challenge for Multimodal Large Language Models (MLLMs), as current R1-based methodologies often prioritize text-centric reasoning deri…
Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking
Qiangqiang Wu, Yi Yu, Chenqi Kong +5
With the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has l…
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
Bo Fang, Wenhao Wu, Qiangqiang Wu +2
Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e…
DistinctAD: Distinctive Audio Description Generation in Contexts
Bo Fang, Wenhao Wu, Qiangqiang Wu +2
Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automa…
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
Jia Wan, Qiangqiang Wu, Wei Lin +1
The existing crowd counting models require extensive training data, which is time-consuming to annotate. To tackle this issue, we propose a simple yet effective crowd counting meth…