5 papers
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
Yue Heng Yeo, Haoyang Li, Yizhou Peng +6
Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augme…
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
Haoyang Li, Changsong Liu, Wei Rao +3
Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce…
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
Haoyang Li, Xuyi Zhuang, Azmat Adnan +6
Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenT…
Aligning Generative Speech Enhancement with Perceptual Feedback
Haoyang Li, Nana Hou, Yuchen Hu +6
Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives tha…
Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
Shreyas Gopal, Ashutosh Anshul, Haoyang Li +3
Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for…