activity
20222025
most citedTowards Frame-level Quality Predictions of Synthetic Speech

3 citations · 4 across the 6 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2025

Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice

Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann +3

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustra…

eess.AS20253 cited

Towards Frame-level Quality Predictions of Synthetic Speech

Michael Kuhlmann, Fritz Seebauer, Petra Wagner +1

While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This…

eess.AS2025

Speech Synthesis along Perceptual Voice Quality Dimensions

Frederik Rautenberg, Michael Kuhlmann, Fritz Seebauer +3

While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we…

eess.AS2023

On Feature Importance and Interpretability of Speaker Representations

Frederik Rautenberg, Michael Kuhlmann, Jana Wiechmann +3

Unsupervised speech disentanglement aims at separating fast varying from slowly varying components of a speech signal. In this contribution, we take a closer look at the embedding…

eess.AS2023

Investigating Speaker Embedding Disentanglement on Natural Read Speech

Michael Kuhlmann, Adrian Meise, Fritz Seebauer +2

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to i…

eess.AS20221 cited

Investigation into Target Speaking Rate Adaptation for Voice Conversion

Michael Kuhlmann, Fritz Seebauer, Janek Ebbers +2

Disentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a…