collaborators

5 papers

cs.SD2025

SLM-TTA: A Framework for Test-Time Adaptation of Generative Spoken Language Models

Yuan-Kuei Wu, Yang Liu, Yiteng Huang +9

Spoken Language Models (SLMs) are increasingly central to modern speech-driven applications, but performance degrades under acoustic shift - real-world noise, reverberation, and mi…

cs.LG2025

Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining

Ruizhe Huang, Kexuan Zhang, Yihao Fang +1

This study investigates small-scale pretraining for Small Language Models (SLMs) to enable efficient use of limited data and compute, improve accessibility in low-resource settings…

cs.CL2025

BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights

Chan-Jan Hsu, Yi-Cheng Lin, Chia-Chun Lin +10

We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyp…

eess.AS2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…

eess.AS2024

Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation

Ruizhe Huang, Mahsa Yarmohammadi, Sanjeev Khudanpur +1

Existing research suggests that automatic speech recognition (ASR) models can benefit from additional contexts (e.g., contact lists, user specified vocabulary). Rare words and name…