activity
20242026
most citedEnhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

1 citations · 1 across the 9 of their papers we have counts for

collaborators

16 papers

cs.SD2026

Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models

Haoqin Sun, Chenyang Lyu, Shiwan Zhao +5

Despite the growing success of Large Speech Language Models (LSLMs) in processing short-term acoustic signals, their extension to long-form audio understanding is severely bottlene…

cs.CL2026

Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue

Yuhang Jia, Pei Liu, Haoqin Sun +6

End-to-end Spoken Language Models (SLMs) hold great potential for paralinguistic perception, and numerous studies have aimed to enhance their capabilities, particularly for empathe…

cs.SD2025

AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation

Hui Wang, Jinghua Zhao, Junyang Cheng +5

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only li…

cs.SD2025

MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning

Haoqin Sun, Chenyang Lyu, Xiangyu Kong +9

Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discr…

cs.SD2025

Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering

Jinghua Zhao, Hang Su, Lichun Fan +4

With the rapid progress of large audio-language models (LALMs), audio question answering (AQA) has emerged as a challenging task requiring both fine-grained audio understanding and…

cs.SD2025

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

Hui Wang, Cheng Liu, Junyang Chen +7

Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generaliza…