2 papers
cs.SD2025
TISDiSS: A Training-Time and Inference-Time Scalable Framework for Discriminative Source Separation
Yongsheng Feng, Yuetonghui Xu, Jiehui Luo +4
Source separation is a fundamental task in speech, music, and audio processing, and it also provides cleaner and larger data for training generative models. However, improving sepa…
cs.SD2024
SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement
Zihao Wang, Le Ma, Yongsheng Feng +3
Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can h…