8 papers
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
Meiqi Cao, Xiangbo Shu, Jiachao Zhang +3
Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current l…
ASTRA: Let Arbitrary Subjects Transform in Video Editing
Fei Shen, Weihao Xu, Rui Yan +4
While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entang…
Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
Zheng Yin, Chengjian Li, Xiangbo Shu +3
Comprehensively and flexibly capturing the complex spatio-temporal dependencies of human motion is critical for multi-person motion prediction. Existing methods grapple with two pr…
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
Ling Xing, Hongyu Qu, Rui Yan +2
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where even…
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
Binqian Xu, Xiangbo Shu, Haiyang Mei +3
Multimodal Large Language Models (MLLMs) have made significant advancements, demonstrating powerful capabilities in processing and understanding multimodal data. Fine-tuning MLLMs…
MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action Recognition
Hongyu Qu, Rui Yan, Xiangbo Shu +3
Recent few-shot action recognition (FSAR) methods typically perform semantic matching on learned discriminative features to achieve promising performance. However, most FSAR method…