papers

Publications (56)

cs.CL2024

Building a Taiwanese Mandarin Spoken Language Model: A First Attempt

Chih-Kai Yang, Yu-Kuan Fu, Chen-An Li +18

This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech…

eess.AS2026

How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation

Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang +13

Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-…

eess.AS2025

Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach

Yi-Cheng Lin, Huang-Cheng Chou, Hung-yi Lee

While subgroup disparities and performance bias are increasingly studied in computational research, fairness in categorical Speech Emotion Recognition (SER) remains underexplored.…

eess.AS2025

Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget

Andy T. Liu, Yi-Cheng Lin, Haibin Wu +2

Despite their impressive success, training foundation models remains computationally costly. This paper investigates how to efficiently train speech foundation models with self-sup…

cs.SD2026

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin +3

Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures i…

eess.AS2024

Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +13

Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The i…