6 papers
High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
Shen Zheng, Jiaran Cai, Yuansheng Guan +7
Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or re…
Mobile-Agent-v3: Fundamental Agents for GUI Automation
Jiabo Ye, Xi Zhang, Haiyang Xu +12
This paper introduces GUI-Owl, a foundational GUI agent model that achieves state-of-the-art performance among open-source end-to-end models on ten GUI benchmarks across desktop an…
Adaptive Duration Model for Text Speech Alignment
Junjie Cao
Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-…
VideoGuard: Protecting Video Content from Unauthorized Editing
Junjie Cao, Kaizhou Li, Xinchun Yu +2
With the rapid development of generative technology, current generative models can generate high-fidelity digital content and edit it in a controlled manner. However, there is a ri…
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
Hongxiang Li, Yaowei Li, Yuhang Yang +4
Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleto…
UVCG: Leveraging Temporal Consistency for Universal Video Protection
KaiZhou Li, Jindong Gu, Xinchun Yu +3
The security risks of AI-driven video editing have garnered significant attention. Although recent studies indicate that adding perturbations to images can protect them from malici…