3 citations · 4 across the 3 of their papers we have counts for
6 papers
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
Frederik Rautenberg, Jana Wiechmann, Petra Wagner +1
We introduce a system capable of faithfully modifying the perceptual voice quality of creak while preserving the speaker's perceived identity. While it is well known that high crea…
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
Michael Kuhlmann, Alexander Werning, Thilo von Neumann +1
A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cann…
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann +3
The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustra…
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
Adrian Meise, Tobias Cord-Landwehr, Reinhold Haeb-Umbach
Diffusion models have been shown to achieve natural-sounding enhancement of speech degraded by noise or reverberation. However, their simultaneous denoising and dereverberation cap…
Towards Frame-level Quality Predictions of Synthetic Speech
Michael Kuhlmann, Fritz Seebauer, Petra Wagner +1
While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This…
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
Tobias Cord-Landwehr, Tobias Gburrek, Marc Deegen +1
We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed s…