6 papers
Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
Ziqian Ning, Shuai Wang, Yuepeng Jiang +5
Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have…
MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
Jixun Yao, Qing Wang, Pengcheng Guo +4
Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information…
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
Ziqian Ning, Shuai Wang, Pengcheng Zhu +4
Streaming voice conversion has become increasingly popular for its potential in real-time applications. The recently proposed DualVC 2 has achieved robust and high-quality streamin…
Accent-VITS:accent transfer for end-to-end TTS
Linhan Ma, Yongmao Zhang, Xinfa Zhu +4
Accent transfer aims to transfer an accent from a source speaker to synthetic speech in the target speaker's voice. The main challenge is how to effectively disentangle speaker tim…
VITS-Based Singing Voice Conversion Leveraging Whisper and multi-scale F0 Modeling
Ziqian Ning, Yuepeng Jiang, Zhichao Wang +2
This paper introduces the T23 team's system submitted to the Singing Voice Conversion Challenge 2023. Following the recognition-synthesis framework, our singing conversion model is…
DualVC: Dual-mode Voice Conversion using Intra-model Knowledge Distillation and Hybrid Predictive Coding
Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu +4
Voice conversion is an increasingly popular technology, and the growing number of real-time applications requires models with streaming conversion capabilities. Unlike typical (non…