1 citations · 1 across the 1 of their papers we have counts for
5 papers
High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
Shen Zheng, Jiaran Cai, Yuansheng Guan +7
Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or re…
Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
Xingpei Ma, Shenneng Huang, Jiaran Cai +5
Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, ex…
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
Xingpei Ma, Jiaran Cai, Yuansheng Guan +3
Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference…
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
Yiqiao Huang, Yuancheng Wang, Jiaqi Li +4
In debating, rebuttal is one of the most critical stages, where a speaker addresses the arguments presented by the opposing side. During this process, the speaker synthesizes their…
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Yuancheng Wang, Haoyue Zhan, Liwei Liu +7
The recent large-scale text-to-speech (TTS) systems are usually grouped as autoregressive and non-autoregressive systems. The autoregressive systems implicitly model duration but e…