3 papers
eess.AS2025
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
Qingliang Meng, Yuqing Deng, Wei Liang +3
Natural and human-like speech depends on the coordination between prosodic timing and acoustic realization: duration modeling shapes rhythmic structure, while waveform generation d…
cs.CL2025
Chinchunmei at SemEval-2025 Task 11: Boosting the Large Language Model's Capability of Emotion Perception using Contrastive Learning
Tian Li, Yujian Sun, Huizhi Liang
The SemEval-2025 Task 11, Bridging the Gap in Text-Based Emotion Detection, introduces an emotion recognition challenge spanning over 28 languages. This competition encourages rese…
cs.CL2025
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems
Qingliang Meng, Pengju Ren, Tian Li +2
Automatic speech recognition (ASR) systems normally consist of an acoustic model (AM) and a language model (LM). The acoustic model estimates the probability distribution of text g…