works on

From the 1 of 18 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu +2

The paper proposes IAAN, a training‑free method that identifies and amplifies specific neurons inside the audio encoder of large audio‑language models to improve recognition of fin…

cs.SD2026

MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo +7

While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capabil…

cs.SD2026

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

Leonardo Haw-Yang Foo, Chih-Kai Yang, Chen-An Li +2

Large Audio-Language Models show consistent performance gains across speech and audio benchmarks, yet high scores may not reflect true auditory perception. If a model can answer qu…

cs.SD2026

Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models

Lok-Lam Ieong, Chia-Chien Chen, Chih-Kai Yang +3

Chain-of-thought (CoT) prompting has been extended to large audio-language models (LALMs) to elicit reasoning, yet enhancing its effectiveness without training remains challenging.…

cs.SD2026

SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models

Chih-Kai Yang, Yen-Ting Piao, Tzu-Wen Hsu +8

Knowledge editing enables targeted updates without retraining, but prior work focuses on textual or visual facts, leaving abstract auditory perceptual knowledge underexplored. We i…

cs.SD2025

Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations

Bo-Han Feng, Chien-Feng Liu, Yu-Hsuan Li Liang +9

Large audio-language models (LALMs) extend text-based LLMs with auditory understanding, offering new opportunities for multimodal applications. While their perception, reasoning, a…