6 papers
Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections
Zeyu Xu, Andreas Brendel, Albert G. Prinn +1
Room impulse responses (RIRs) are fundamental to audio data augmentation, acoustic signal processing, and immersive audio rendering. While geometric simulators such as the image so…
Robust Online Overdetermined Independent Vector Analysis Based on Bilinear Decomposition
Kang Chen, Xianrui Wang, Yichen Yang +6
Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (Over…
DeePAQ: A Perceptual Audio Quality Metric Based On Foundational Models and Weakly Supervised Learning
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado +1
This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the…
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
Spatial target speaker extraction isolates a desired speaker's voice in multi-speaker environments using spatial information, such as the direction of arrival (DoA). Although recen…
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
Generative speech enhancement methods based on generative adversarial networks (GANs) and diffusion models have shown promising results in various speech enhancement tasks. However…
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoG…