6 papers · 1 filter
Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations
Lyonel Behringer, Andreas Brendel
High-dimensional representations of pretrained speech foundation models have proven beneficial for objective speech quality and intelligibility prediction. While existing work on n…
SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
Nagashree K. S. Rao, Shrishti Saha Shetu, Mohamed Elminshawi +2
Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion…
A Comparison of Generative and Discriminative Methods for Speech Enhancement: Robustness, Complexity, and Hallucination
Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
In this study, we conduct a comprehensive comparative analysis of generative and discriminative deep learning-based speech enhancement methods, specifically in noise reduction task…
UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension
Kishan Gupta, Srikanth Korse, Andreas Brendel +2
In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs…
On Improving Error Resilience of Neural End-to-End Speech Coders
Kishan Gupta, Nicola Pia, Srikanth Korse +3
Error resilient tools like Packet Loss Concealment (PLC) and Forward Error Correction (FEC) are essential to maintain a reliable speech communication for applications like Voice ov…
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
Andreas Brendel, Nicola Pia, Kishan Gupta +3
Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end tra…