6 citations · 6 across the 12 of their papers we have counts for
5 papers · 1 filter
Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1
Few-step diffusion and flow-matching text-to-speech (TTS) models are usually trained with local objectives, such as conditional flow matching, reconstruction, and stop prediction.…
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3
Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…
Context-Aware ASR for Mandarin Technical Lectures
Ho-Lam Chung, Yiming Chen, Hung-yi Lee
Technical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) t…
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
Ho-Lam Chung, Yiming Chen, Hung-yi Lee
Neural audio codecs are widely used as tokenizers for spoken language models, but they are optimized for waveform reconstruction rather than autoregressive prediction. This mismatc…
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
Hao-Hui Xie, Ho-Lam Chung, Yi-Cheng Lin +4
Large Audio-Language Models (LALMs) typically struggle with localized dialectal prosody due to the scarcity of specialized corpora. We present TW-Sound580K, a Taiwanese audio-text…