5 citations · 7 across the 6 of their papers we have counts for
4 papers · 1 filter
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis
Guanghui Xu, Wei Song, Zhengchen Zhang +3
Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which ma…
Emotion recognition by fusing time synchronous and time asynchronous representations
Wen Wu, Chao Zhang, Philip C. Woodland
In this paper, a novel two-branch neural network model structure is proposed for multimodal emotion recognition, which consists of a time synchronous branch (TSB) and a time asynch…
Combination of Deep Speaker Embeddings for Diarisation
Guangzhi Sun, Chao Zhang, Phil Woodland
Significant progress has recently been made in speaker diarisation after the introduction of d-vectors as speaker embeddings extracted from neural network (NN) speaker classifiers…
Neural Kalman Filtering for Speech Enhancement
Wei Xue, Gang Quan, Chao Zhang +3
Statistical signal processing based speech enhancement methods adopt expert knowledge to design the statistical models and linear filters, which is complementary to the deep neural…