10 papers
SLM-TTA: A Framework for Test-Time Adaptation of Generative Spoken Language Models
Yuan-Kuei Wu, Yang Liu, Yiteng Huang +9
Spoken Language Models (SLMs) are increasingly central to modern speech-driven applications, but performance degrades under acoustic shift - real-world noise, reverberation, and mi…
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
Yufeng Yang, Yiteng Huang, Yong Xu +9
With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-co…
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
Yang Liu, Li Wan, Yiteng Huang +5
Smart glasses are increasingly positioned as the next-generation interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in re…
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang +8
Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech L…
Directional Source Separation for Robust Speech Recognition on Smart Glasses
Tiantian Feng, Ju Lin, Yiteng Huang +7
Modern smart glasses leverage advanced audio sensing and machine learning technologies to offer real-time transcribing and captioning services, considerably enriching human experie…
Effective Integration of KAN for Keyword Spotting
Anfeng Xu, Biqiao Zhang, Shuyu Kong +4
Keyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (…