5 papers
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
Yuang Zheng, Dongxu Chen, Yuxiang Mei +3
Large-scale multilingual ASR (mASR) models such as Whisper achieve strong performance but incur high computational and latency costs, limiting their deployment on resource-constrai…
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
Shuangyuan Chen, Shuang Wei, Dongxing Xu +1
To enhance the performance of end-to-end (E2E) speech recognition systems in noisy or low signal-to-noise ratio (SNR) conditions, this paper introduces NoisyD-CT, a novel tri-stage…
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
Hanfang Cui, Longfei Song, Li Li +2
Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematical…
SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge
Yuxiang Mei, Yuang Zheng, Dongxing Xu +1
This paper describes SHNU multilingual conversational speech recognition system (SHNU-mASR, team name-"maybe"), submitted to Track 1 of the INTERSPEECH 2025 MLC-SLM Challenge. Our…
ICSD: An Open-source Dataset for Infant Cry and Snoring Detection
Qingyu Liu, Longfei Song, Dongxing Xu +1
The detection and analysis of infant cry and snoring events are crucial tasks within the field of audio signal processing. While existing datasets for general sound event detection…