6 papers
Robust Speech Activity Detection in the Presence of Singing Voice
Philipp Grundhuber, Mhd Modar Halimeh, Martin Strauß +1
Speech Activity Detection (SAD) systems often misclassify singing as speech, leading to degraded performance in applications such as dialogue enhancement and automatic speech recog…
Neural Directional Filtering Using a Compact Microphone Array
Weilong Huang, Srikanth Raj Chetupalli, Mhd Modar Halimeh +2
Beamforming with desired directivity patterns using compact microphone arrays is essential in many audio applications. Directivity patterns achievable using traditional beamformers…
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
Philipp Grundhuber, Mhd Modar Halimeh, Emanuël A. P. Habets
This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic…
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
Generative speech enhancement methods based on generative adversarial networks (GANs) and diffusion models have shown promising results in various speech enhancement tasks. However…
Navigating PESQ: Up-to-Date Versions and Open Implementations
Matteo Torcoli, Mhd Modar Halimeh, Emanuël A. P. Habets
Perceptual Evaluation of Speech Quality (PESQ) is an objective quality measure that remains widely used despite its withdrawal by the International Telecommunication Union (ITU). P…
Expanding and Analyzing ODAQ -- the Open Dataset of Audio Quality
Sascha Dick, Christoph Thompson, Chih-Wei Wu +4
The Open Dataset of Audio Quality (ODAQ) was recently introduced to address the scarcity of openly available audio datasets with corresponding subjective quality scores. The datase…