1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content
Sheng Wu, Xiaobao Wang, Longbiao Wang +2
Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, dis…
cs.CL2024
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
Yuchun Shu, Bo Hu, Yifeng He +3
Accurately finding the wrong words in the automatic speech recognition (ASR) hypothesis and recovering them well-founded is the goal of speech error correction. In this paper, we p…
cs.CL2024
An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
Cheng Gong, Erica Cooper, Xin Wang +9
Self-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks. Despite advancements, language…