4 papers
SVoice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
Ziqian Wang, Xianjun Xia, Chuanzeng Huang +1
We present SVoice, the winning system of the Singing Voice Conversion Challenge (SVCC) 2025 for both the in-domain and zero-shot singing style conversion tracks. Built on the s…
LDCodec: A high quality neural audio codec with low-complexity decoder
Jiawei Jiang, Linping Xu, Dejun Zhang +3
Neural audio coding has been shown to outperform classical audio coding at extremely low bitrates. However, the practical application of neural audio codecs is still limited by the…
UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
Ziqian Wang, Zikai Liu, Yike Zhu +6
Generative modeling has recently achieved remarkable success across image, video, and audio domains, demonstrating powerful capabilities for unified representation learning. Yet sp…
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
Ziqian Wang, Xianjun Xia, Xinfa Zhu +1
The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a compr…