13 papers
EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
Junyu Wang, Siyuan Zhang, Peiyuan Jiang +11
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to…
Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection
Jinyu Li, Xiao Wei, Bin Wen +5
Spontaneous speech is a vital non-invasive biomarker for Alzheimer's Disease (AD), yet many systems overlook non-linear structural disruptions and clinical heterogeneity in patholo…
EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning
Siyuan Zhang, Jian Zong, Junyu Wang +7
While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with c…
CECOR: Correction-oriented synthetic data construction for factual error correction
Lei Zhu, Xiaobao Wang, Jianbiao Yang +4
Factual Error Correction (FEC) aims to revise inaccurate text into statements that are factually consistent with external evidence. Although recent methods perform well on single-h…
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang, Ziyang Ma, Yizhou Peng +26
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its con…
MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates
Zikang Huang, Meng Ge, Tianrui Wang +4
Self-supervised learning (SSL) has advanced speech processing. However, existing speech SSL methods typically assume a single sampling rate and struggle with mixed-rate data due to…