collaborators

6 papers

cs.CV2025

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

Shen Zheng, Jiaran Cai, Yuansheng Guan +7

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or re…

cs.AI2025

Mobile-Agent-v3: Fundamental Agents for GUI Automation

Jiabo Ye, Xi Zhang, Haiyang Xu +12

This paper introduces GUI-Owl, a foundational GUI agent model that achieves state-of-the-art performance among open-source end-to-end models on ten GUI benchmarks across desktop an…

cs.SD2025

Adaptive Duration Model for Text Speech Alignment

Junjie Cao

Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-…

cs.CV2025

VideoGuard: Protecting Video Content from Unauthorized Editing

Junjie Cao, Kaizhou Li, Xinchun Yu +2

With the rapid development of generative technology, current generative models can generate high-fidelity digital content and edit it in a controlled manner. However, there is a ri…

cs.CV2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

Hongxiang Li, Yaowei Li, Yuhang Yang +4

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleto…

cs.CV2024

UVCG: Leveraging Temporal Consistency for Universal Video Protection

KaiZhou Li, Jindong Gu, Xinchun Yu +3

The security risks of AI-driven video editing have garnered significant attention. Although recent studies indicate that adding perturbations to images can protect them from malici…