collaborators

7 papers

cs.CV2025

MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis

Shitong Shao, Hongwei Yi, Hanzhong Guo +5

Recently, open-source video diffusion models (VDMs), such as WanX, Magic141 and HunyuanVideo, have been scaled to over 10 billion parameters. These large-scale VDMs have demonstrat…

cs.CV2025

MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Hongwei Yi, Tian Ye, Shitong Shao +10

We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse c…

cs.CV2024

Real-time One-Step Diffusion-based Expressive Portrait Videos Generation

Hanzhong Guo, Hongwei Yi, Daquan Zhou +3

Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. Ho…

cs.CV2024

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

Longtao Zheng, Yifan Zhang, Hanzhong Guo +6

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, ma…

cs.CV2024

Try-On-Adapter: A Simple and Flexible Try-On Paradigm

Hanzhong Guo, Jianfeng Zhang, Cheng Zou +6

Image-based virtual try-on, widely used in online shopping, aims to generate images of a naturally dressed person conditioned on certain garments, providing significant research an…

cs.CR2024

FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent Space

Yiyang Guo, Ruizhe Li, Mude Hui +5

Invisible watermarking is essential for safeguarding digital content, enabling copyright protection and content authentication. However, existing watermarking methods fall short in…