20 papers
A Fast Solver for Interpolating Stochastic Differential Equation Diffusion Models for Speech Restoration
Bunlong Lay, Timo Gerkmann
Diffusion Probabilistic Models (DPMs) are a well-established class of diffusion models for unconditional image generation, while SGMSE+ is a well-established conditional diffusion…
Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation
Rostislav Makarov, Timo Gerkmann
Classifier guidance is a way to control diffusion generation by using a noise-conditioned classifier to steer the sampling process toward a target class. One drawback of classifier…
Your U-Net Dereverberation Model is Secretly an RIR Encoder
Sina Khanagha, Timo Gerkmann
In this work, we analyze the ability of NCSN++ U-Net based audio dereverberation models to capture global room characteristics in their intermediate representations. Through an emp…
Position-Blind Ptychography: Viability of image reconstruction via data-driven variational inference
Simon Welker, Lorenz Kuger, Tim Roith +4
In this work, we present and investigate the novel blind inverse problem of position-blind ptychography, i.e., ptychographic phase retrieval without any knowledge of scan positions…
Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
Danilo de Oliveira, Tal Peer, Timo Gerkmann
Speech enhancement (SE) systems are typically evaluated using a variety of instrumental metrics. The use of automatic speech recognition (ASR) systems to evaluate SE performance is…
Real-Time Streamable Generative Speech Restoration with Flow Matching
Simon Welker, Bunlong Lay, Maris Hillemann +2
Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their…