7 papers
It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement
Zahra Benslimane, Pierre Chouteau, Martyna Poreba +4
Neural network-based multichannel speech enhancement systems achieve strong enhancement performance, but their computational and memory requirements limit deployment on resource-co…
RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices
Z. Benslimane, P. Chouteau, M. Poreba +4
Real-time binaural speech enhancement is constrained by latency, computational cost, and inter-device communication, yet existing efficient solutions predominantly address single-c…
Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement
Nasser-Eddine Monir, Paul Magron, Romain Serizel
Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spec…
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Colombe Mboungou, Mostafa Sadeghi, Jean-Eudes Ayilo +1
Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE,…
Diffusion-based Frameworks for Unsupervised Speech Enhancement
Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel +1
This paper addresses unsupervised diffusion-based single-channel speech enhancement (SE). Prior work in this direction combines a score-based diffusion model trained on clean speec…
Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel +1
We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using no…