4 papers
HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
Xiang Wang, Zhifei Zhang, He Zhang +11
Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanc…
Replace Anyone in Videos
Xiang Wang, Shiwei Zhang, Haonan Qiu +7
The field of controllable human-centric video generation has witnessed remarkable progress, particularly with the advent of diffusion models. However, achieving precise and localiz…
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
Xiang Wang, Shiwei Zhang, Longxiang Tang +4
This report presents UniAnimate-DiT, an advanced project that leverages the cutting-edge and powerful capabilities of the open-source Wan2.1 model for consistent human image animat…
Taming Consistency Distillation for Accelerated Human Image Animation
Xiang Wang, Shiwei Zhang, Hangjie Yuan +5
Recent advancements in human image animation have been propelled by video diffusion models, yet their reliance on numerous iterative denoising steps results in high inference costs…