128 citations · 219 across the 57 of their papers we have counts for
8 papers · 1 filter
DeepA: A Deep Neural Analyzer For Speech And Singing Vocoding
Sergey Nikonorov, Berrak Sisman, Mingyang Zhang +1
Conventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under…
StrengthNet: Deep Learning-based Emotion Strength Assessment for Emotional Speech Synthesis
Rui Liu, Berrak Sisman, Haizhou Li
Recently, emotional speech synthesis has achieved remarkable performance. The emotion strength of synthesized speech can be controlled flexibly using a strength descriptor, which i…
Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion
Zongyang Du, Berrak Sisman, Kun Zhou +1
Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of spe…
VisualTTS: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over
Junchen Lu, Berrak Sisman, Rui Liu +2
In this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis,…
Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style Transfer
Zongyang Du, Berrak Sisman, Kun Zhou +1
Traditional voice conversion(VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in…
Emotional Voice Conversion: Theory, Databases and ESD
Kun Zhou, Berrak Sisman, Rui Liu +1
In this paper, we first provide a review of the state-of-the-art emotional voice conversion research, and the existing emotional speech databases. We then motivate the development…