6 citations · 11 across the 22 of their papers we have counts for
6 papers · 1 filter
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
Jianing Yang, Sheng Li, Takahiro Shinozaki +2
Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details…
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
Zhao Ren, Rathi Adarshi Rammohan, Kevin Scheck +2
Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, a…
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
Haowei Lou, Hye-young Paik, Sheng Li +2
Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging…
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
Jiliang Hu, Zuchao Li, Mengjia Shen +3
Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved gr…
Fusion of Self-supervised Learned Models for MOS Prediction
Zhengdong Yang, Wangjin Zhou, Chenhui Chu +4
We participated in the mean opinion score (MOS) prediction challenge, 2022. This challenge aims to predict MOS scores of synthetic speech on two tracks, the main track and a more c…
Cross-scale Attention Model for Acoustic Event Classification
Xugang Lu, Peng Shen, Sheng Li +2
A major advantage of a deep convolutional neural network (CNN) is that the focused receptive field size is increased by stacking multiple convolutional layers. Accordingly, the mod…