papers

Publications (10)

cs.SD2025

Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis

Sho Inoue, Kun Zhou, Shuai Wang +1

We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel mu…

cs.SD2025

PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs

Sho Inoue, Shai Wang, Haizhou Li

Despite significant progress in neural spoken dialog systems, personality-aware conversation agents -- capable of adapting behavior based on personalities -- remain underexplored d…

cs.SD2025

MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion

Sho Inoue, Shuai Wang, Wanxing Wang +3

In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, w…

cs.CV2021

Style-Restricted GAN: Multi-Modal Translation with Style Restriction Using Generative Adversarial Networks

Sho Inoue, Tad Gonsalves

Unpaired image-to-image translation using Generative Adversarial Networks (GAN) is successful in converting images among multiple domains. Moreover, recent studies have shown a way…

eess.AS2024

Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Zhijun Liu, Shuai Wang, Sho Inoue +2

Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete s…

cs.SD2024

Fine-Grained Quantitative Emotion Editing for Speech Generation

Sho Inoue, Kun Zhou, Shuai Wang +1

It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulati…