4 papers
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
Shenghui Lu, Hukai Huang, Jinanglong Yao +3
This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins…
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
Peijie Chen, Wenhao Guan, Kaidi Wang +4
Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech t…
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
Kaidi Wang, Wenhao Guan, Ziyue Jiang +5
Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speak…
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
Hukai Huang, Jiayan Lin, Kaidi Wang +4
Due to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes…