6 papers
Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics
Jakob Kienegger, Tal Peer, Sina Khanagha +1
Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers…
Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
Danilo de Oliveira, Tal Peer, Timo Gerkmann
Speech enhancement (SE) systems are typically evaluated using a variety of instrumental metrics. The use of automatic speech recognition (ASR) systems to evaluate SE performance is…
Real-Time Streamable Generative Speech Restoration with Flow Matching
Simon Welker, Bunlong Lay, Maris Hillemann +2
Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their…
Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
Danilo de Oliveira, Tal Peer, Jonas Rochdi +1
Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets…
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
Julius Richter, Danilo de Oliveira, Tal Peer +1
We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach…
Real-Time Streaming Mel Vocoding with Generative Flow Matching
Simon Welker, Tal Peer, Timo Gerkmann
The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on gen…