most citedTowards Frame-level Quality Predictions of Synthetic Speech

3 citations · 4 across the 3 of their papers we have counts for

collaborators

6 papers

eess.AS2026

Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis

Frederik Rautenberg, Jana Wiechmann, Petra Wagner +1

We introduce a system capable of faithfully modifying the perceptual voice quality of creak while preserving the speaker's perceived identity. While it is well known that high crea…

eess.AS20261 cited

Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts

Michael Kuhlmann, Alexander Werning, Thilo von Neumann +1

A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cann…

eess.AS2025

Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice

Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann +3

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustra…

eess.AS2025

On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation

Adrian Meise, Tobias Cord-Landwehr, Reinhold Haeb-Umbach

Diffusion models have been shown to achieve natural-sounding enhancement of speech degraded by noise or reverberation. However, their simultaneous denoising and dereverberation cap…

eess.AS20253 cited

Towards Frame-level Quality Predictions of Synthetic Speech

Michael Kuhlmann, Fritz Seebauer, Petra Wagner +1

While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This…

eess.AS2025

Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering

Tobias Cord-Landwehr, Tobias Gburrek, Marc Deegen +1

We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed s…