6 papers
Preference Optimization for Non-Verbal Vocalization Synthesis
Haoyang Li, Chenglin Xu, Junchuan Zhao +4
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains po…
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
Yue Heng Yeo, Haoyang Li, Yizhou Peng +6
Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augme…
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
Haoyang Li, Changsong Liu, Wei Rao +3
Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce…
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
Haoyang Li, Xuyi Zhuang, Azmat Adnan +6
Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenT…
Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
Shreyas Gopal, Ashutosh Anshul, Haoyang Li +3
Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for…
Aligning Generative Speech Enhancement with Perceptual Feedback
Haoyang Li, Nana Hou, Yuchen Hu +6
Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives tha…