3 papers
cs.CL2026
UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition
Ying Fang, Xiaofei Li
This paper proposes a unimodal aggregation (UMA) based nonautoregressive model for both English and Mandarin speech recognition. The original UMA explicitly segments and aggregates…
eess.AS2025
VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification
Pengyu Wang, Ying Fang, Xiaofei Li
Reverberant speech, denoting the speech signal degraded by reverberation, contains crucial knowledge of both anechoic source speech and room impulse response (RIR). This work propo…
eess.AS2024
Mamba for Streaming ASR Combined with Unimodal Aggregation
Ying Fang, Xiaofei Li
This paper works on streaming automatic speech recognition (ASR). Mamba, a recently proposed state space model, has demonstrated the ability to match or surpass Transformers in var…