29 citations · 31 across the 6 of their papers we have counts for
6 papers
Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study
Darwin Jelestin Muthu, Navya Gupta, Wei Lin Tay +3
Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored.…
Multi-Speaker Multi-Style Speech Synthesis with Timbre and Style Disentanglement
Wei Song, Yanghao Yue, Ya-jie Zhang +3
Disentanglement of a speaker's timbre and style is very important for style transfer in multi-speaker multi-style text-to-speech (TTS) scenarios. With the disentanglement of timbre…
Singing Voice Synthesis with Vibrato Modeling and Latent Energy Representation
Yingjie Song, Wei Song, Wei Zhang +4
This paper proposes an expressive singing voice synthesis system by introducing explicit vibrato modeling and latent energy representation. Vibrato is essential to the naturalness…
ViDA-MAN: Visual Dialog with Digital Humans
Tong Shen, Jiawei Zuo, Fan Shi +7
We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text o…
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis
Guanghui Xu, Wei Song, Zhengchen Zhang +3
Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which ma…
Telerobotic Pointing Gestures Shape Human Spatial Cognition
John-John Cabibihan, Wing-Chee So, Sujin Saj +1
This paper aimed to explore whether human beings can understand gestures produced by telepresence robots. If it were the case, they can derive meaning conveyed in telerobotic gestu…