5 papers
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
Binhe Yu, Zhen Wang, Kexin Li +6
Multi-subject customization aims to synthesize multiple user-specified subjects into a coherent image. To address issues such as subjects missing or conflicts, recent works incorpo…
CoMo: Compositional Motion Customization for Text-to-Video Generation
Youcan Xu, Zhen Wang, Jiaxin Shi +6
While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods f…
Zero-shot Compositional Action Recognition with Neural Logic Constraints
Gefan Ye, Lin Li, Kexin Li +2
Zero-shot compositional action recognition (ZS-CAR) aims to identify unseen verb-object compositions in the videos by exploiting the learned knowledge of verb and object primitives…
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
Kexin Li, Tao Jiang, Zongxin Yang +3
Interactive Video Object Segmentation (iVOS) is a challenging task that requires real-time human-computer interaction. To improve the user experience, it is important to consider t…
Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual Segmentation
Kexin Li, Zongxin Yang, Yi Yang +1
Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects that accurately align with the corresponding audio. However, existing methods of…