From the 1 of 12 linked papers with an AI index.
12 papers
Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1
The paper introduces a Speech Representation Fréchet Distance loss (SR‑FD) that regularizes few‑step diffusion/flow‑matching TTS models by matching the statistics of Whisper and CT…
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2
While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation q…
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
Ho Lam Chung, Bo-Xuan Zheng, Cheng-Chieh Huang +8
Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment common Taiwanese text, and their…
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3
Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…
Context-Aware ASR for Mandarin Technical Lectures
Ho-Lam Chung, Yiming Chen, Hung-yi Lee
Technical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) t…
Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR
Ho Lam Chung, Yiming Chen, Dau-Cheng Lyu +2
End-to-end ASR models transcribe in a single pass, leaving no room for the decoder to revisit hard inputs. We propose LatentASR, a parameter-efficient method that adds continuous l…