2 papers
eess.AS2025
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
Xinfa Zhu, Wenjie Tian, Xinsheng Wang +4
Text-to-Audio (TTA) generation is an emerging area within AI-generated content (AIGC), where audio is created from natural language descriptions. Despite growing interest, developi…
eess.AS2025
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
Xinfa Zhu, Lei He, Yujia Xiao +4
Style voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker's identity. However, previous style voice conve…