1.5k citations · 1.9k across the 18 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2020
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis
Guanghui Xu, Wei Song, Zhengchen Zhang +3
Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which ma…
eess.AS2019★ 17 cited
Towards adversarial learning of speaker-invariant representation for speech emotion recognition
Ming Tu, Yun Tang, Jing Huang +2
Speech emotion recognition (SER) has attracted great attention in recent years due to the high demand for emotionally intelligent speech interfaces. Deriving speaker-invariant repr…