8 papers
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
Meiqi Cao, Xiangbo Shu, Jiachao Zhang +3
Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current l…
ASTRA: Let Arbitrary Subjects Transform in Video Editing
Fei Shen, Weihao Xu, Rui Yan +4
While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entang…
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
Hongyu Qu, Xiangbo Shu, Rui Yan +3
Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coar…
Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
Zheng Yin, Chengjian Li, Xiangbo Shu +3
Comprehensively and flexibly capturing the complex spatio-temporal dependencies of human motion is critical for multi-person motion prediction. Existing methods grapple with two pr…
Vision-centric Token Compression in Large Language Model
Ling Xing, Alex Jinpeng Wang, Rui Yan +2
Real-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to trillions of parameters. This dua…
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
Ling Xing, Hongyu Qu, Rui Yan +2
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where even…