collaborators

6 papers

eess.AS2026

Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics

Jakob Kienegger, Tal Peer, Sina Khanagha +1

Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers…

eess.AS2026

Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement

Danilo de Oliveira, Tal Peer, Timo Gerkmann

Speech enhancement (SE) systems are typically evaluated using a variety of instrumental metrics. The use of automatic speech recognition (ASR) systems to evaluate SE performance is…

eess.SP2026

Real-Time Streamable Generative Speech Restoration with Flow Matching

Simon Welker, Bunlong Lay, Maris Hillemann +2

Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their…

eess.AS2025

Are These Even Words? Quantifying the Gibberishness of Generative Speech Models

Danilo de Oliveira, Tal Peer, Jonas Rochdi +1

Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets…

eess.AS2025

LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

Julius Richter, Danilo de Oliveira, Tal Peer +1

We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach…

eess.AS2025

Real-Time Streaming Mel Vocoding with Generative Flow Matching

Simon Welker, Tal Peer, Timo Gerkmann

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on gen…