6 citations · 19 across the 40 of their papers we have counts for
8 papers · 1 filter
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3
Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…
How Contrastive Decoding Enhances Large Audio Language Models
Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin +1
While Contrastive Decoding (CD) has been proposed to enhance Large Audio Language Models (LALMs), it has not been evaluated at scale, and the underlying mechanisms driving its succ…
Latent-Mark: An Audio Watermark Robust to Neural Codec Compression
Yen-Shan Chen, Shih-Yu Lai, Ying-Jung Tsou +5
While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, they remain vulnerable to neural compressi…
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
Hao-Hui Xie, Ho-Lam Chung, Yi-Cheng Lin +4
Large Audio-Language Models (LALMs) typically struggle with localized dialectal prosody due to the scarcity of specialized corpora. We present TW-Sound580K, a Taiwanese audio-text…
How Does Instrumental Music Help SingFake Detection?
Xuanjun Chen, Chia-Yu Hu, I-Ming Lin +8
Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how inst…
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +3
Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically pr…