activity
20242026
collaborators

11 papers

cs.SD2026

What Makes Synthetic Speech Sound Sarcastic? A Prosody-Controlled Perception Study

Zhu Li, Shekhar Nayak, Matt Coler

Prosody plays an important role in sarcasm perception, yet previous studies have relied on naturally produced speech that lacks fine-grained control over individual acoustic dimens…

cs.CL2026

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

Reihaneh Amooie, Yun Hao, Wietse de Vries +3

This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographi…

cs.CL2026

Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework

Zhu Li, Yuqing Zhang, Xiyuan Gao +2

Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However,…

cs.CL2026

Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

Zhu Li, Yuqing Zhang, Xiyuan Gao +2

Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often re…

cs.MM2026

SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning

Zhu Li, Yongjian Chen, Huiyuan Lai +3

Multimodal sarcasm detection requires resolving pragmatic incongruity across textual, acoustic, and visual cues through cross-modal reasoning. To enable robust sarcasm reasoning wi…

cs.CL2025

Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

Zhu Li, Xiyuan Gao, Yuqing Zhang +2

Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior…