6 papers
ERA-Solver: Error-Robust Adams Solver for Fast Sampling of Diffusion Probabilistic Models
Shengming Li, Luping Liu, Runnan Li +1
Though denoising diffusion probabilistic models (DDPMs) have achieved remarkable generation results, the low sampling efficiency of DDPMs still limits further applications. Since D…
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
Chenpeng Du, Qi Chen, Tianyu He +5
While recent research has made significant progress in speech-driven talking face generation, the quality of the generated video still lags behind that of real recordings. One reas…
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Dongchao Yang, Jinchuan Tian, Xu Tan +9
Large Language models (LLM) have demonstrated the capability to handle a variety of generative tasks. This paper presents the UniAudio system, which, unlike prior task-specific app…
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
Anni Tang, Tianyu He, Xu Tan +2
Talking face generation aims at generating photo-realistic video portraits of a target person driven by input audio. Due to its nature of one-to-many mapping from the input audio t…
VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer
Liyang Chen, Zhiyong Wu, Runnan Li +4
Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotono…
Empowering Diffusion Models on the Embedding Space for Text Generation
Zhujin Gao, Junliang Guo, Xu Tan +4
Diffusion models have achieved state-of-the-art synthesis quality on both visual and audio tasks, and recent works further adapt them to textual data by diffusing on the embedding…