1 citations · 1 across the 5 of their papers we have counts for
5 papers
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang, Ning Cheng +2
Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech…
EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang, Jianzong Wang +2
There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited…
SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech Model
Jianzong Wang, Xulong Zhang, Haobin Tang +3
In recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there…
Dynamic Alignment Mask CTC: Improved Mask-CTC with Aligned Cross Entropy
Xulong Zhang, Haobin Tang, Jianzong Wang +3
Because of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autor…
QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang, Jianzong Wang +2
Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TT…