collaborators

5 papers

cs.CV2026

Contact Matrix: Enhancing Dance Motion Synthesis with Precise Interaction Modeling

Xuhai Chen, Zhi Cen, Huaijin Pi +3

Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints and a limited feasible soluti…

cs.CV2026

Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs

Jiazheng Xing, Chao Xu, Hangjie Yuan +4

Multimodal Large Language Models (MLLMs) have propelled the field of few-shot action recognition (FSAR). However, preliminary explorations in this area primarily focus on generatin…

cs.CV2026

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

Jiazheng Xing, Fei Du, Hangjie Yuan +7

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and…

cs.CV2026

SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training

Mengmeng Wang, Dengyang Jiang, Liuzhuozheng Li +6

Denoising-based diffusion transformers, despite their strong generation performance, suffer from inefficient training convergence. Existing methods addressing this issue, such as R…

cs.CV2024

Visual Object Tracking across Diverse Data Modalities: A Review

Mengmeng Wang, Teli Ma, Shuo Xin +5

Visual Object Tracking (VOT) is an attractive and significant research area in computer vision, which aims to recognize and track specific targets in video sequences where the targ…