4 papers
Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
Danilo de Oliveira, Tal Peer, Jonas Rochdi +1
Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets…
Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
Navin Raj Prabhu, Danilo de Oliveira, Nale Lehmann-Willenbrock +1
Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have…
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
Julius Richter, Danilo de Oliveira, Tal Peer +1
We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach…
Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture
Julius Richter, Danilo de Oliveira, Timo Gerkmann
We study diffusion-based speech enhancement using a Schrodinger bridge formulation and extend the EDM2 framework to this setting. We employ time-dependent preconditioning of networ…