collaborators

6 papers

cs.CV2026

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration

Linrui Tian, Qi Wang, Bang Zhang

Real-time streaming joint audio-video generation for character animation requires a generator to speak the requested transcript, maintain visual identity across chunks, and run wit…

cs.CV2026

SyncAnyone: Implicit Disentanglement via Progressive Self-Correction for Lip-Syncing in the wild

Xindi Zhang, Dechao Meng, Steven Xiao +3

High-quality AI-powered video dubbing demands precise audio-lip synchronization, high-fidelity visual generation, and faithful preservation of identity and background. Most existin…

cs.CV2025

Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation

Steven Xiao, Xindi Zhang, Dechao Meng +3

Real-time portrait animation is essential for interactive applications such as virtual assistants and live avatars, requiring high visual fidelity, temporal coherence, ultra-low la…

cs.CV2025

Wan-S2V: Audio-Driven Cinematic Video Generation

Xin Gao, Li Hu, Siqi Hu +20

Current state-of-the-art (SOTA) methods for audio-driven character animation demonstrate promising performance for scenarios primarily involving speech and singing. However, they o…

cs.CV2025

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

Dechao Meng, Steven Xiao, Xindi Zhang +5

Audio-driven portrait animation, which synthesizes realistic videos from reference images using audio signals, faces significant challenges in real-time generation of high-fidelity…

cs.CV2025

EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Linrui Tian, Siqi Hu, Qi Wang +2

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing meth…