53 citations · 239 across the 41 of their papers we have counts for
5 papers · 1 filter
Integrating Uncertainty into Neural Network-based Speech Enhancement
Huajian Fang, Dennis Becker, Stefan Wermter +1
Supervised masking approaches in the time-frequency domain aim to employ deep neural networks to estimate a multiplicative mask to extract clean speech. This leads to a single esti…
Partially Adaptive Multichannel Joint Reduction of Ego-noise and Environmental Noise
Huajian Fang, Niklas Wittmer, Johannes Twiefel +2
Human-robot interaction relies on a noise-robust audio processing module capable of estimating target speech from audio recordings impacted by environmental noise, as well as self-…
Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement
Huajian Fang, Tal Peer, Stefan Wermter +1
Speech enhancement in the time-frequency domain is often performed by estimating a multiplicative mask to extract clean speech. However, most neural network-based methods perform p…
Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder
Huajian Fang, Guillaume Carbajal, Stefan Wermter +1
Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the train…
Multimodal Target Speech Separation with Voice and Face References
Leyuan Qu, Cornelius Weber, Stefan Wermter
Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the…