activity
20212024
most citedMake-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

47 citations · 159 across the 19 of their papers we have counts for

collaborators

19 papers

cs.SD2024

MulliVC: Multi-lingual Voice Conversion With Cycle Consistency

Jiawei Huang, Chen Zhang, Yi Ren +6

Voice conversion aims to modify the source speaker's voice to resemble the target speaker while preserving the original speech content. Despite notable advancements in voice conver…

cs.CV20242 cited

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Zhenhui Ye, Tianyun Zhong, Yi Ren +11

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait vid…

cs.CV2023

C2G2: Controllable Co-speech Gesture Generation with Latent Diffusion Model

Longbin Ji, Pengfei Wei, Yi Ren +3

Co-speech gesture generation is crucial for automatic digital avatar animation. However, existing methods suffer from issues such as unstable training and temporal inconsistency, p…

cs.CV2023

Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis

Zhenhui Ye, Ziyue Jiang, Yi Ren +5

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbit…

eess.AS202316 cited

Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Ziyue Jiang, Yi Ren, Zhenhui Ye +9

Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. Howeve…

cs.SD202310 cited

Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Jiawei Huang, Yi Ren, Rongjie Huang +7

Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consist…