collaborators

6 papers

cs.LG2025

ERA-Solver: Error-Robust Adams Solver for Fast Sampling of Diffusion Probabilistic Models

Shengming Li, Luping Liu, Runnan Li +1

Though denoising diffusion probabilistic models (DDPMs) have achieved remarkable generation results, the low sampling efficiency of DDPMs still limits further applications. Since D…

cs.CV2025

DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder

Chenpeng Du, Qi Chen, Tianyu He +5

While recent research has made significant progress in speech-driven talking face generation, the quality of the generated video still lags behind that of real recordings. One reas…

cs.SD2024

UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Dongchao Yang, Jinchuan Tian, Xu Tan +9

Large Language models (LLM) have demonstrated the capability to handle a variety of generative tasks. This paper presents the UniAudio system, which, unlike prior task-specific app…

cs.CV2024

Memories are One-to-Many Mapping Alleviators in Talking Face Generation

Anni Tang, Tianyu He, Xu Tan +2

Talking face generation aims at generating photo-realistic video portraits of a target person driven by input audio. Due to its nature of one-to-many mapping from the input audio t…

cs.CV2024

VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer

Liyang Chen, Zhiyong Wu, Runnan Li +4

Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotono…

cs.CL2024

Empowering Diffusion Models on the Embedding Space for Text Generation

Zhujin Gao, Junliang Guo, Xu Tan +4

Diffusion models have achieved state-of-the-art synthesis quality on both visual and audio tasks, and recent works further adapt them to textual data by diffusing on the embedding…