5 papers
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
Zhu Li, Xiyuan Gao, Yuqing Zhang +2
Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior…
Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
Xiyuan Gao, Shekhar Nayak, Matt Coler
Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance…
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
Ömer Tarik Özyilmaz, Matt Coler, Matias Valdenegro-Toro
Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic (MSA), they struggle with dialectal speech. We investigate the effect of fine-t…
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
Reihaneh Amooie, Wietse de Vries, Yun Hao +3
Automatic Speech Recognition (ASR) performance for low-resource languages is still far behind that of higher-resource languages such as English, due to a lack of sufficient labeled…
AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation
Xiyuan Gao, Shubhi Bansal, Kushaan Gowda +4
Detecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in…