most citedEmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

Haobin Tang, Xulong Zhang, Ning Cheng +2

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech…

cs.SD20231 cited

EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis

Haobin Tang, Xulong Zhang, Jianzong Wang +2

There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited…

cs.SD2023

SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech Model

Jianzong Wang, Xulong Zhang, Haobin Tang +3

In recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there…

cs.SD2023

Dynamic Alignment Mask CTC: Improved Mask-CTC with Aligned Cross Entropy

Xulong Zhang, Haobin Tang, Jianzong Wang +3

Because of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autor…

cs.SD2023

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

Haobin Tang, Xulong Zhang, Jianzong Wang +2

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TT…