works on

From the 1 of 44 linked papers with an AI index.

activity
20242026
most citedOn the social bias of speech self-supervised models

6 citations · 6 across the 20 of their papers we have counts for

collaborators
Showing cs.SDShow all

10 papers · 1 filter

cs.SD2026

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech

Ho Lam Chung, Bo-Xuan Zheng, Cheng-Chieh Huang +8

Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment common Taiwanese text, and their…

cs.SD2026

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

Ho Lam Chung, Yiming Chen, Dau-Cheng Lyu +2

End-to-end ASR models transcribe in a single pass, leaving no room for the decoder to revisit hard inputs. We propose LatentASR, a parameter-efficient method that adds continuous l…

cs.SD2026

Latent-Mark: An Audio Watermark Robust to Neural Codec Compression

Yen-Shan Chen, Shih-Yu Lai, Ying-Jung Tsou +5

While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, they remain vulnerable to neural compressi…

cs.SD2026

Membership Inference Attacks against Large Audio Language Models

Jia-Kai Dong, Yu-Xiang Lin, Hung-Yi Lee

We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs. Using Multi-modal Blind Baselines based on textual, spectral and prosodic features, we demons…

cs.SD2026

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

Hao-Hui Xie, Ho-Lam Chung, Yi-Cheng Lin +4

Large Audio-Language Models (LALMs) typically struggle with localized dialectal prosody due to the scarcity of specialized corpora. We present TW-Sound580K, a Taiwanese audio-text…

cs.SD2026

LLM-Codec: Neural Audio Codec Meets Language Model Objectives

Ho-Lam Chung, Yiming Chen, Hung-yi Lee

Neural audio codecs are widely used as tokenizers for spoken language models, but they are optimized for waveform reconstruction rather than autoregressive prediction. This mismatc…