collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen +18

We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes M…

cs.CL2026

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham +3

Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, w…

cs.CL2025

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Yiming Gao, Bin Wang, Chengwei Wei +2

Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after al…

cs.CL2025

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

Wenyu Zhang, Yingxu He, Geyu Lin +9

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues s…

cs.CL2025

Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data

Qiongqiong Wang, Hardik Bhupendra Sailor, Tianchi Liu +7

Recent speech-LLMs have shown impressive performance in tasks like transcription and translation, yet they remain limited in understanding the paralinguistic aspects of speech cruc…

cs.CL2025

Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models

Qiongqiong Wang, Hardik B. Sailor, Jeremy H. M. Wong +6

Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextu…