3 papers
eess.AS2024
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
Ruibo Fu, Xin Qi, Zhengqi Wen +10
Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media…
cs.SD2024
ADD 2022: the First Audio Deep Synthesis Detection Challenge
Jiangyan Yi, Ruibo Fu, Jianhua Tao +17
Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios.…
eess.AS2024
MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
Ruibo Fu, Shuchen Shi, Hongming Guo +12
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements…