1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang, Jianzong Wang +2
There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited…
SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech Model
Jianzong Wang, Xulong Zhang, Haobin Tang +3
In recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there…
Dynamic Alignment Mask CTC: Improved Mask-CTC with Aligned Cross Entropy
Xulong Zhang, Haobin Tang, Jianzong Wang +3
Because of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autor…
QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang, Jianzong Wang +2
Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TT…