6 papers
Localizing and Editing Knowledge in Large Audio-Language Models
Sung Kyun Chung, Jiaheng Dong, Qiuchi Hu +3
Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual information. Yet they are trained…
ImKWS: Test-Time Adaptation for Keyword Spotting with Class Imbalance
Hanyu Ding, Yang Xiao, Jiaheng Dong +1
Keyword spotting (KWS) identifies words for voice assistants, but environmental noise frequently reduces accuracy. Standard adaptation fixes this issue and strictly requires origin…
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
Hongyu Jin, Siyi Wang, Yang Xiao +10
Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory. Despite recent progress in large…
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
Xiaofeng Yu, Jiaheng Dong, Jean Honorio +3
Speech emotion recognition plays an important role in various applications. However, most existing approaches predict a single emotion label, oversimplifying the inherently ambiguo…
E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
Jiaheng Dong, Hong Jia, Soumyajit Chatterjee +3
Speech Foundation Models encounter significant performance degradation when deployed in real-world scenarios involving acoustic domain shifts, such as background noise and speaker…
Test-Time Adaptation for Speech Emotion Recognition
Jiaheng Dong, Hong Jia, Ting Dang
The practical utility of Speech Emotion Recognition (SER) systems is undermined by their fragility to domain shifts, such as speaker variability, the distinction between acted and…