works on

From the 2 of 21 linked papers with an AI index.

activity
20242026
collaborators

21 papers

cs.CL2026

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu +2

The paper investigates drift in model-generated timestamps for autoregressive ASR systems and introduces REDDIT, a replay‑based distribution editing post‑training method that corre…

cs.SD2026

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu +2

The paper proposes IAAN, a training‑free method that identifies and amplifies specific neurons inside the audio encoder of large audio‑language models to improve recognition of fin…

cs.SD2026

MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo +7

While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capabil…

eess.AS2026

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

Ke-Han Lu, Keqi Deng, Ruchao Fan +2

Speech large language models (Speech LLMs) typically encode speech into sequences far longer than text, creating a major efficiency bottleneck during autoregressive decoding. A com…

cs.SD2026

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3

Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…

eess.AS2026

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models

Chun-Wei Chen, Tzu-Quan Lin, Ke-Han Lu +2

Speech Language Models achieve reasoning capabilities, but are often hindered by massive parameter counts and a tendency to prioritize linguistic priors over acoustic features. Whi…