5 papers · 1 filter
A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement
Jie Zhang, Haoyin Yan, Xiaofei Li
It is promising to design a single model that can suppress various distortions and improve speech quality, i.e., universal speech enhancement (USE). Compared to supervised learning…
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
Wang Dai, Xiaofei Li, Archontis Politis +1
In end-to-end multi-channel speech enhancement, the traditional approach of designating one microphone signal as the reference for processing may not always yield optimal results.…
RVAE-EM: Generative speech dereverberation based on recurrent variational auto-encoder and convolutive transfer function
Pengyu Wang, Xiaofei Li
In indoor scenes, reverberation is a crucial factor in degrading the perceived quality and intelligibility of speech. In this work, we propose a generative dereverberation method.…
Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors
Di Liang, Nian Shao, Xiaofei Li
This work proposes a frame-wise online/streaming end-to-end neural diarization (FS-EEND) method in a frame-in-frame-out fashion. To frame-wisely detect a flexible number of speaker…
FN-SSL: Full-Band and Narrow-Band Fusion for Sound Source Localization
Yabo Wang, Bing Yang, Xiaofei Li
Extracting direct-path spatial features is critical for sound source localization in adverse acoustic environments. This paper proposes a full-band and narrow-band fusion network f…