collaborators

16 papers

eess.AS2026

Unsupervised Speech Recognition at the Syllable Level

Liming Wang, Kai-Wei Chang, Kunio Kashino +3

Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in…

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

cs.CL2026

TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild

Kai-Wei Chang, Yi-Cheng Lin, Huang-Cheng Chou +9

Speech technologies have advanced rapidly and serve diverse populations worldwide. However, many languages remain underrepresented due to limited resources. In this paper, we intro…

cs.CL2026

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering

Cheng-Kuang Chang, Kai-Wei Chang, Alexander H. Liu +1

Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the internal mechanism by which they coor…

eess.AS2026

RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection

Wenze Ren, Ke-Han Lu, Kai-Wei Chang +8

Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) e…

cs.CL2026

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

Chan-Jan Hsu, Liang-Hsuan Tseng, Yi-Cheng Lin +5

Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving attributes like speaker and emotion, se…