4 papers
Fish Audio S2 Technical Report
Shijia Liao, Yuxuan Wang, Songting Liu +11
We introduce Fish Audio S2, an open-sourced text-to-speech system featuring multi-speaker, multi-turn generation, and, most importantly, instruction-following control via natural-l…
Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models
Nikita Kuzmin, Songting Liu, Kong Aik Lee +1
Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural a…
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
Haoyang Li, Yuchen Hu, Chen Chen +3
Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures neede…
Zero-shot Voice Conversion with Diffusion Transformers
Songting Liu
Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre…