3 papers
cs.SD2025
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
Bingsong Bai, Yizhong Geng, Fengping Wang +4
Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing meth…
eess.AS2025
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
Bingsong Bai, Qihang Lu, Wenbing Yang +8
Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets,…
cs.SD2024
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
Hongming Guo, Ruibo Fu, Yizhong Geng +9
Text-to-audio (TTA) model is capable of generating diverse audio from textual prompts. However, most mainstream TTA models, which predominantly rely on Mel-spectrograms, still face…