12 citations · 27 across the 9 of their papers we have counts for
22 papers
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
Zongyang Du, Junchen Lu, Kun Zhou +2
Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling fo…
EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
Perry Lam, Huayun Zhang, Nancy F. Chen +1
Neural models are known to be over-parameterized, and recent work has shown that sparse text-to-speech (TTS) models can outperform dense models. Although a plethora of sparse metho…
Controllable Accented Text-to-Speech Synthesis
Rui Liu, Berrak Sisman, Guanglai Gao +1
Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1). Accented TTS synthesis is challenging as L2 is diffe…
DeepA: A Deep Neural Analyzer For Speech And Singing Vocoding
Sergey Nikonorov, Berrak Sisman, Mingyang Zhang +1
Conventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under…
StrengthNet: Deep Learning-based Emotion Strength Assessment for Emotional Speech Synthesis
Rui Liu, Berrak Sisman, Haizhou Li
Recently, emotional speech synthesis has achieved remarkable performance. The emotion strength of synthesized speech can be controlled flexibly using a strength descriptor, which i…
Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style Transfer
Zongyang Du, Berrak Sisman, Kun Zhou +1
Traditional voice conversion(VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in…