7 papers
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
Bo Fang, Yuxin Song, Qiangqiang Wu +3
Complex video reasoning remains a significant challenge for Multimodal Large Language Models (MLLMs), as current R1-based methodologies often prioritize text-centric reasoning deri…
Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking
Qiangqiang Wu, Yi Yu, Chenqi Kong +5
With the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has l…
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
Bo Fang, Wenhao Wu, Qiangqiang Wu +2
Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e…
Density-based Object Detection in Crowded Scenes
Chenyang Zhao, Jia Wan, Antoni B. Chan
Compared with the generic scenes, crowded scenes contain highly-overlapped instances, which result in: 1) more ambiguous anchors during training of object detectors, and 2) more pr…
Video Individual Counting for Moving Drones
Yaowu Fan, Jia Wan, Tao Han +2
Video Individual Counting (VIC) has received increasing attention for its importance in intelligent video surveillance. Existing works are limited in two aspects, i.e., dataset and…
Embodied Crowd Counting
Runling Long, Yunlong Wang, Jia Wan +5
Occlusion is one of the fundamental challenges in crowd counting. In the community, various data-driven approaches have been developed to address this issue, yet their effectivenes…