11 papers
FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates
Jiaqi Li, Chaoren Wang, Xiaohai Tian +9
Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying informati…
Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention
Ming-Hao Hsu, Xiaohai Tian, Jun Zhang +1
Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We reveal that this gap is not a uniform cognitive deficit. Evaluating two architect…
End-to-end Listen, Look, Speak and Act
Siyin Wang, Wenyi Yu, Xianzhao Chen +4
Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and interruptions. Realizing these capabi…
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
Ming-Hao Hsu, Xueyao Zhang, Xiaohai Tian +2
Recent advancements in Large Speech-Language Models have significantly bridged the gap between acoustic signals and linguistic understanding. However, a persistent performance disp…
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
Zhixian Zhao, Wenjie Tian, Lei Xie
Multimodal emotion analysis is shifting from static classification to generative reasoning. Beyond simple label prediction, robust affective reasoning must synthesize fine-grained…
Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
Siyin Wang, Zengrui Jin, Changli Tang +26
In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities…