4 papers · 1 filter
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
Frederik Rautenberg, Jana Wiechmann, Petra Wagner +1
We introduce a system capable of faithfully modifying the perceptual voice quality of creak while preserving the speaker's perceived identity. While it is well known that high crea…
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann +3
The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustra…
Speech Synthesis along Perceptual Voice Quality Dimensions
Frederik Rautenberg, Michael Kuhlmann, Fritz Seebauer +3
While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we…
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
Yuying Xie, Michael Kuhlmann, Frederik Rautenberg +2
Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important…