6 papers
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
Qibing Bai, Sho Inoue, Shuai Wang +3
Accent normalization converts foreign-accented speech into native-like speech while preserving speaker identity. We propose a novel pipeline using self-supervised discrete tokens a…
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel mu…
Hierarchical Control of Emotion Rendering in Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains chal…
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
Sho Inoue, Shai Wang, Haizhou Li
Despite significant progress in neural spoken dialog systems, personality-aware conversation agents -- capable of adapting behavior based on personalities -- remain underexplored d…
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
Sho Inoue, Shuai Wang, Wanxing Wang +3
In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, w…
Fine-Grained Quantitative Emotion Editing for Speech Generation
Sho Inoue, Kun Zhou, Shuai Wang +1
It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulati…