7 papers
Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
Danilo de Oliveira, Tal Peer, Timo Gerkmann
Speech enhancement (SE) systems are typically evaluated using a variety of instrumental metrics. The use of automatic speech recognition (ASR) systems to evaluate SE performance is…
Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture
Julius Richter, Danilo de Oliveira, Timo Gerkmann
We study diffusion-based speech enhancement using a Schrodinger bridge formulation and extend the EDM2 framework to this setting. We employ time-dependent preconditioning of networ…
Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
Danilo de Oliveira, Tal Peer, Jonas Rochdi +1
Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets…
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
Julius Richter, Danilo de Oliveira, Tal Peer +1
We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach…
Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
Navin Raj Prabhu, Danilo de Oliveira, Nale Lehmann-Willenbrock +1
Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have…
Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
Danilo de Oliveira, Julius Richter, Jean-Marie Lemercier +2
Diffusion models have found great success in generating high quality, natural samples of speech, but their potential for density estimation for speech has so far remained largely u…