6 papers · 1 filter
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3
Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…
Context-Aware ASR for Mandarin Technical Lectures
Ho-Lam Chung, Yiming Chen, Hung-yi Lee
Technical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) t…
Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR
Ho Lam Chung, Yiming Chen, Dau-Cheng Lyu +2
End-to-end ASR models transcribe in a single pass, leaving no room for the decoder to revisit hard inputs. We propose LatentASR, a parameter-efficient method that adds continuous l…
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
Feiyu Zhao, Yiming Chen, Wenhuan Lu +3
Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models generate responses that are s…
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
Ho-Lam Chung, Yiming Chen, Hung-yi Lee
Neural audio codecs are widely used as tokenizers for spoken language models, but they are optimized for waveform reconstruction rather than autoregressive prediction. This mismatc…
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
Yiming Chen, Xianghu Yue, Xiaoxue Gao +4
Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primaril…