activity
20242026
collaborators

5 papers

cs.CL2026

UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition

Ying Fang, Xiaofei Li

This paper proposes a unimodal aggregation (UMA) based nonautoregressive model for both English and Mandarin speech recognition. The original UMA explicitly segments and aggregates…

eess.AS2025

VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification

Pengyu Wang, Ying Fang, Xiaofei Li

Reverberant speech, denoting the speech signal degraded by reverberation, contains crucial knowledge of both anechoic source speech and room impulse response (RIR). This work propo…

eess.AS2025

CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR

Nian Shao, Rui Zhou, Pengyu Wang +4

In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) p…

eess.AS2024

Mamba for Streaming ASR Combined with Unimodal Aggregation

Ying Fang, Xiaofei Li

This paper works on streaming automatic speech recognition (ASR). Mamba, a recently proposed state space model, has demonstrated the ability to match or surpass Transformers in var…

cs.SD2024

RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization

Bing Yang, Changsheng Quan, Yabo Wang +7

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffu…