activity
20242026
most citedTowards audio language modeling -- an overview

6 citations · 19 across the 40 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2026

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3

Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…

cs.SD2026

How Contrastive Decoding Enhances Large Audio Language Models

Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin +1

While Contrastive Decoding (CD) has been proposed to enhance Large Audio Language Models (LALMs), it has not been evaluated at scale, and the underlying mechanisms driving its succ…

cs.SD2026

Latent-Mark: An Audio Watermark Robust to Neural Codec Compression

Yen-Shan Chen, Shih-Yu Lai, Ying-Jung Tsou +5

While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, they remain vulnerable to neural compressi…

cs.SD2026

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

Hao-Hui Xie, Ho-Lam Chung, Yi-Cheng Lin +4

Large Audio-Language Models (LALMs) typically struggle with localized dialectal prosody due to the scarcity of specialized corpora. We present TW-Sound580K, a Taiwanese audio-text…

cs.SD2025

How Does Instrumental Music Help SingFake Detection?

Xuanjun Chen, Chia-Yu Hu, I-Ming Lin +8

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how inst…

cs.SD2025

ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +3

Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically pr…