5 papers · 1 filter
SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
Nagashree K. S. Rao, Shrishti Saha Shetu, Mohamed Elminshawi +2
Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion…
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
Srikanth Korse, Mohamed Elminshawi, Emanuel A. P. Habets +1
The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment a…
Dynamic Slimmable Networks for Efficient Speech Separation
Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets
Recent progress in speech separation has been largely driven by advances in deep neural networks, yet their high computational and memory requirements hinder deployment on resource…
Beamformer-Guided Target Speaker Extraction
Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets
We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of…
Noise-Robust Adaptation Control for Supervised Acoustic System Identification Exploiting A Noise Dictionary
Thomas Haubner, Andreas Brendel, Mohamed Elminshawi +1
We present a noise-robust adaptation control strategy for block-online supervised acoustic system identification by exploiting a noise dictionary. The proposed algorithm takes adva…