6 papers · 1 filter
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel mu…
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
Sho Inoue, Shai Wang, Haizhou Li
Despite significant progress in neural spoken dialog systems, personality-aware conversation agents -- capable of adapting behavior based on personalities -- remain underexplored d…
Hierarchical Control of Emotion Rendering in Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains chal…
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
Sho Inoue, Shuai Wang, Wanxing Wang +3
In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, w…
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representa…
Fine-Grained Quantitative Emotion Editing for Speech Generation
Sho Inoue, Kun Zhou, Shuai Wang +1
It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulati…