4 papers
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
David Sasu, Natalie Schluter
We show the performance of Automatic Speech Recognition (ASR) systems that use semi-supervised speech representations can be boosted by a complimentary pitch accent detection modul…
Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition
Ziwei Gong, Pengyuan Shi, Kaan Donbekci +5
Speech Emotion Recognition (SER) has seen significant progress with deep learning, yet remains challenging for Low-Resource Languages (LRLs) due to the scarcity of annotated data.…
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody
David Sasu, Kweku Andoh Yamoah, Benedict Quartey +1
Enabling robots to accurately interpret and execute spoken language instructions is essential for effective human-robot collaboration. Traditional methods rely on speech recognitio…
Akan Cinematic Emotions (ACE): A Multimodal Multi-party Dataset for Emotion Recognition in Movie Dialogues
David Sasu, Zehui Wu, Ziwei Gong +5
In this paper, we introduce the Akan Conversation Emotion (ACE) dataset, the first multimodal emotion dialogue dataset for an African language, addressing the significant lack of r…