4 papers
Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study
Zhen Wang, TianRui Wu, RongQi Han +3
Synthetic speech offers scalable supervision for automatic speech recognition (ASR), but its benefit depends on text selection, reference speech, and augmentation scale. We present…
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
Hao Wu, RongQi Han, Zhen Wang +2
This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker dia…
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
Qingliang Meng, Yuqing Deng, Wei Liang +3
Natural and human-like speech depends on the coordination between prosodic timing and acoustic realization: duration modeling shapes rhythmic structure, while waveform generation d…
ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition
Qingliang Meng, Hao Wu, Wei Liang +2
The deep integration of large language models and automatic speech recognition systems has become a promising research direction with high practical value. To address the overfitti…