9 papers
DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
Yuepeng Jiang, Huakang Chen, Ziqian Ning +7
Generating full-length, high-quality songs is challenging, as it requires maintaining long-term coherence both across text and music modalities and within the music modality itself…
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
Jixun Yao, Yuguang Yang, Yu Pan +5
Integrating human feedback to align text-to-speech (TTS) system outputs with human preferences has proven to be an effective approach for enhancing the robustness of language model…
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
Guobin Ma, Jixun Yao, Ziqian Ning +4
Zero-shot voice conversion (VC) aims to transfer timbre from a source speaker to any unseen target speaker while preserving linguistic content. Growing application scenarios demand…
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
Zhao Guo, Ziqian Ning, Guobin Ma +1
Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming…
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
Yuepeng Jiang, Ziqian Ning, Shuai Wang +5
In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ens…
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
Huakang Chen, Yuepeng Jiang, Guobin Ma +7
Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in…