4 papers
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
Ming-Hao Hsu, Xueyao Zhang, Xiaohai Tian +2
Recent advancements in Large Speech-Language Models have significantly bridged the gap between acoustic signals and linguistic understanding. However, a persistent performance disp…
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang +8
Multilingual speech foundation models such as Whisper are trained on web-scale data, where data for each language consists of a myriad of regional varieties. However, different reg…
TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
Ming-Hao Hsu, Liang-Hsuan Tseng, Hung-yi Lee +1
We propose Text-Aligned Speech Tokens with Multiple Layer-Aggregation (TASLA), which is a text-aligned speech tokenization framework that aims to address the problem that under a l…
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
Ming-Hao Hsu, Hung-yi Lee
Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due…