activity
20242026
collaborators

8 papers

cs.SD2026

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3

Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…

cs.SD2026

Context-Aware ASR for Mandarin Technical Lectures

Ho-Lam Chung, Yiming Chen, Hung-yi Lee

Technical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) t…

cs.SD2026

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

Ho Lam Chung, Yiming Chen, Dau-Cheng Lyu +2

End-to-end ASR models transcribe in a single pass, leaving no room for the decoder to revisit hard inputs. We propose LatentASR, a parameter-efficient method that adds continuous l…

cs.SD2026

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

Feiyu Zhao, Yiming Chen, Wenhuan Lu +3

Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models generate responses that are s…

cs.SD2026

LLM-Codec: Neural Audio Codec Meets Language Model Objectives

Ho-Lam Chung, Yiming Chen, Hung-yi Lee

Neural audio codecs are widely used as tokenizers for spoken language models, but they are optimized for waveform reconstruction rather than autoregressive prediction. This mismatc…

eess.AS2026

MORE: Multi-Objective Adversarial Attacks on Speech Recognition

Xiaoxue Gao, Zexin Li, Yiming Chen +1

The emergence of large-scale automatic speech recognition (ASR) models such as Whisper has greatly expanded their adoption across diverse real-world applications. Ensuring robustne…