collaborators

11 papers

cs.CL2026

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

Xuanru Zhou, Yiwen Shao, Jiahong Li +1

Multimodal large language models (MLLMs) are typically built through a multi-stage pipeline consisting of cross-modal alignment, supervised fine-tuning (SFT), and preference optimi…

eess.AS2026

Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech

Xuanru Zhou, Jiachen Lian, Henry Hong +2

Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the cont…

eess.AS2026

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo +3

Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-langu…

cs.SD2026

Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

Xuanru Zhou, Yiwen Shao, Wei-Cheng Tseng +1

Current audio pre-training seeks to learn unified representations for broad audio understanding tasks, but it remains fragmented and is fundamentally bottlenecked by its reliance o…

cs.CL2026

K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function

Shuhe Li, Chenxu Guo, Jiachen Lian +18

Evaluating young children's language is challenging for automatic speech recognizers due to high-pitched voices, prolonged sounds, and limited data. We introduce K-Function, a fram…

eess.AS2025

LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness

Zongli Ye, Jiachen Lian, Akshaj Gupta +18

Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used…