10 papers
SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
Xiaoyu Yang, Xuenan Xu, Wenyi Yu +10
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder…
FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations
Haolong Zheng, Siyin Wang, Xulin Fan +2
Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large langu…
End-to-end Listen, Look, Speak and Act
Siyin Wang, Wenyi Yu, Xianzhao Chen +4
Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and interruptions. Realizing these capabi…
Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
Yudong Yang, Xuezhen Zhang, Zhifeng Han +6
Recent progress in LLMs has enabled understanding of audio signals, but has also exposed new safety risks arising from complex audio inputs that are inadequately handled by current…
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
Kaimeng Jia, Minzhu Tu, Zengrui Jin +2
Dysarthria is a speech disorder characterized by impaired intelligibility and reduced communicative effectiveness. Automatic dysarthria assessment provides a scalable, cost-effecti…
Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
Siyin Wang, Zengrui Jin, Changli Tang +26
In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities…